Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This action is in response to the original filing on 07/17/2024. Claims 21-40 are pending and have been considered below.
Information Disclosure Statement
3. The information disclosure statement (IDS(s)) submitted on 09/19/2024, 02/19/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Double Patenting
4. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 4, 8-13, 15-19 of U.S. Patent No. US 12,079,726 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following mapping below. Each corresponding limitation is either identical or does not have a patentable, nonobvious distinction unless otherwise noted.
Instant Application 18/775,643
Patent No.: US 12,079,726 B2
Claim 21
Claims 8, 11, 13
Claim 22
Claim 8
Claim 23
Claim 8
Claim 24
Claim 10
Claim 25
Claim 12
Claim 26
Claim 13
Claim 27
Claim 9
Claim 28
Claims 15, 17
Claim 29
Claim 15
Claim 30
Claim 15
Claim 31
Claim 18
Claim 32
Claim 19
Claim 33
Claim 16
Claim 34
Claims 1, 4
Claim 35
Claim 1
Claim 36
Claim 1
Claim 37
Claim 3
Claim 38
Claim 5
Claim 39
Claim 6
Claim 40
Claim 2
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-5, 9-13, 18 of U.S. Patent No. US 11,604,992 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following mapping below. Each corresponding limitation is either identical or does not have a patentable, nonobvious distinction unless otherwise noted.
Instant Application 18/775,643
Patent No.: US 11,604,992 B2
Claim 21
Claims 1, 3
Claim 22
Claim 9
Claim 23
Claim 10
Claim 24
Claim 11
Claim 25
Claim 12
Claim 26
Claim 13
Claim 27
Claim 1
Claim 28
Claims 1, 2, 3
Claim 29
Claim 9
Claim 30
Claim 10
Claim 31
Claim 12
Claim 32
Claim 13
Claim 33
Claim 18
Claim 34
Claims 1, 3
Claim 35
Claim 1
Claim 36
Claim 1
Claim 37
Claim 2
Claim 38
Claim 4
Claim 39
Claim 5
Claim 40
Claim 1
Claim Rejections - 35 USC § 102
5. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
6. Claims 21-25, 27-38, and 40 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Zhou et al. (U.S. Patent Application Pub. No. US 20190354837 A1).
Claim 21: Zhou teaches a method for generating a neural network architecture (i.e. the policy network uses (215) that network embedding to automatically generate adaptations to the neural network architecture configuration. In one or more embodiments, the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048]), the method comprising:
sampling training data from a training data store (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]);
determining, based on an initial probability distribution, a sample neural network architecture (i.e. the new layers may be generated in a stochastic way. FIG. 6 depicts a methodology to facilitate architecture configuration exploration using probability mass functions. Hence, in one or more embodiments, a goal of the insert LSTM is to define the probability mass function (p.m.f.) (e.g., PMFP 350-P) to sample (650) the features of the new layer to be generated. For each feature, mapping of the LSTM state output to the p.m.f. may be done by a lookup table 345. For example, if there are three (3) candidate values for the feature of convolution width, the LSTM state output determines three (3) probability values corresponding to them; para. [0062]) from a neural network architecture space (i.e. actions of scale and insert are mapped to a search space to define the neural network architectures … layer-by-layer search aims to find the optimal architecture with a search granularity of predefined layers, the neural network architecture is defined (705) by stacking these layers, an LSTM in the policy network chooses (710) the layer type and the corresponding hyperparameters (e.g., filter width); para. [0066, 0068]);
performing an evaluation of the sampled training data using the sample neural network architecture (i.e. the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048, 0076, 0103]);
updating, based on the evaluation, the initial probability distribution to generate an updated probability distribution (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0103]);
determining whether termination criteria is satisfied (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077]); and
when it is determined that the termination criteria is satisfied, generating a result neural network architecture based on the updated probability distribution (i.e. the updated policy network is used for the next episode … At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0077, 0103, 0105]).
Claim 22: Zhou teaches the method of claim 21. Zhou further teaches wherein performing the evaluation of the sampled training data (i.e. the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048, 0076, 0103]) comprises: evaluating the sampled training data using the sampled neural network architecture to compute a gradient of a loss function associated with the sampled neural network architecture (i.e. The ADAM optimization algorithm was used for training each KWS model, a cross entropy (CE) loss function was used for training … The parameters θ of the policy network may be optimized via gradient descent; para. [0088, 0089, 0115]).
Claim 23: Zhou teaches the method of claim 22. Zhou further teaches wherein updating the initial probability distribution to generate the updated probability distribution comprises: updating the initial probability distribution to generate the updated probability distribution based on the computed gradient of the loss function (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0077, 0103]).
Claim 24: Zhou teaches the method of claim 21. Zhou further teaches wherein determining whether the termination criteria is satisfied comprises comparing a first accuracy of a neural network architecture associated with the initial probability distribution (i.e. Given a set of sample networks, performance curves for each network may be obtained. For each network xi, a validation accuracy ai and training time ti may be obtained; para. [0076, 0087, 0094]) and a second accuracy of a neural network architecture associated with the updated probability distribution (i.e. Each model was evaluated after training and an action is selected according to the current policy in order to transform the network. At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0062, 0077, 0103]) based on a predetermined threshold (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077, 0098]).
Claim 25: Zhou teaches the method of claim 21. Zhou further teaches comprising: training the result neural network architecture (i.e. The top eight models from each episode were progressively selected as baseline models to the next episode. We train the best models for longer training time to get SOTA performance; para. [0096]) using training data from the training data store (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]).
Claim 27: Zhou teaches the method of claim 21. Zhou further teaches comprising: when it is not determined that the termination criteria is satisfied, performing another iteration to further update the probability distribution (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0076, 0077]).
Claim 28: Zhou teaches a method for generating a neural network architecture (i.e. the policy network uses (215) that network embedding to automatically generate adaptations to the neural network architecture configuration. In one or more embodiments, the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048]), the method comprising:
determining, based on an initial probability distribution, a sample neural network architecture (i.e. the new layers may be generated in a stochastic way. FIG. 6 depicts a methodology to facilitate architecture configuration exploration using probability mass functions. Hence, in one or more embodiments, a goal of the insert LSTM is to define the probability mass function (p.m.f.) (e.g., PMFP 350-P) to sample (650) the features of the new layer to be generated. For each feature, mapping of the LSTM state output to the p.m.f. may be done by a lookup table 345. For example, if there are three (3) candidate values for the feature of convolution width, the LSTM state output determines three (3) probability values corresponding to them; para. [0062]) from a neural network architecture space (i.e. actions of scale and insert are mapped to a search space to define the neural network architectures … layer-by-layer search aims to find the optimal architecture with a search granularity of predefined layers, the neural network architecture is defined (705) by stacking these layers, an LSTM in the policy network chooses (710) the layer type and the corresponding hyperparameters (e.g., filter width); para. [0066, 0068]);
updating, based on an evaluation of sampled training data using the sample neural network architecture (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]), the initial probability distribution to generate an updated probability distribution (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0103]);
determining whether termination criteria is satisfied by comparing a first accuracy of a neural network architecture associated with the initial probability distribution (i.e. Given a set of sample networks, performance curves for each network may be obtained. For each network xi, a validation accuracy ai and training time ti may be obtained; para. [0076, 0087, 0094]) and a second accuracy of a neural network architecture associated with the updated probability distribution (i.e. Each model was evaluated after training and an action is selected according to the current policy in order to transform the network. At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0062, 0077, 0103]) based on a predetermined threshold (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077, 0098]); and
when it is determined that the termination criteria is satisfied, generating a result neural network architecture based on the updated probability distribution (i.e. the updated policy network is used for the next episode … At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0077, 0103, 0105]).
Claims 29-38 and 40 are similar in scope to Claims 21-25, 27 and are rejected under a similar rationale.
Claim Rejections – 35 USC § 103
7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
8. Claims 26 and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Vasudevan et al. (U.S. Patent Application Pub. No. US 20190026639 A1).
Claim 26: Zhou teaches the method of claim 21. Zhou further teaches wherein the initial probability distribution is determined based on a surrogate neural network architecture (i.e. a search can start with any baseline models, a well-designed or even a rudimentary one … a policy network embodiment 300 uses a network embedding 320 to represent the input neural network configuration 302; para. [0058, 0059, 0061, 0062]).
Zhou does not explicitly teach was determined using a dataset surrogate, wherein the dataset surrogate comprises a different set of training data from the training data store than a set of training data used when sampling training data from a training data store.
However, Vasudevan teaches a surrogate neural network architecture that was determined using a dataset surrogate (i.e. the system can effectively determine the architecture of the convolutional cells on a smaller data set and then re-use the same cell architecture across a range of data and computational scales … The neural architecture search system 100 is a system that obtains training data 102 for training a convolutional neural network to perform a particular task and a validation set 104 for evaluating the performance of the convolutional neural network on the particular task and uses the training data 102 and the validation set 104 to determine an network architecture for a child CNN that is configured to perform the image processing task … Once trained values of the controller parameters have been determined, i.e., once the training of the controller neural network 110 has satisfied some termination criteria, the system determines a final architecture for the first convolutional cell (and any other convolutional cells that are defined by the output sequences generated by the controller neural network); para. [0006, 0016, 0035]), wherein the dataset surrogate comprises a different set of training data from the training data store than a set of training data used when sampling training data from a training data store (i.e. the system can effectively determine the architecture of the convolutional cells on a smaller data set and then re-use the same cell architecture across a range of data and computational scales. In particular, the system can effectively employ the resulting learned architecture to perform image processing tasks with reduced computational budgets that match or outperform streamlined architectures targeted to mobile and embedded platforms; para. [0006, 0038]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Zhou to include the feature of Vasudevan. One would have been motivated to make this modification because the architecture learned using a smaller dataset can be reused across different data scales, thereby reducing the computational resources.
Claim 39 is similar in scope to Claim 26 and is rejected under a similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Sun et al. (Pub. No. US 20200082275 A1), Disclosed are a neural network architecture search apparatus and method and a computer readable recording medium. The neural network architecture search method comprises: defining a search space used as a set of architecture parameters describing the neural network architecture; performing sampling on the architecture parameters in the search space based on parameters of a control unit, to generate at least one sub-neural network architecture; performing training on each sub-neural network architecture by minimizing a loss function including an inter-class loss and a center loss.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TAN H TRAN/Primary Examiner, Art Unit 2141