Prosecution Insights
Last updated: October 02, 2026
Application No. 18/775,643

PROBABILISTIC NEURAL NETWORK ARCHITECTURE GENERATION

Non-Final OA §102§103§DOUBLEPATENT
Filed
Jul 17, 2024
Priority
Nov 02, 2018 — continuation of 11/604,992 +1 more
Examiner
TRAN, TAN H
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
195 granted / 320 resolved
+0.9% vs TC avg
Strong +33% interview lift
Without
With
+32.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
46 currently pending
Career history
374
Total Applications
across all art units

Statute-Specific Performance

§101
13.4%
-26.6% vs TC avg
§103
59.8%
+19.8% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
6.3%
-33.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 320 resolved cases

Office Action

§102 §103 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 2. This action is in response to the original filing on 07/17/2024. Claims 21-40 are pending and have been considered below. Information Disclosure Statement 3. The information disclosure statement (IDS(s)) submitted on 09/19/2024, 02/19/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Double Patenting 4. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 4, 8-13, 15-19 of U.S. Patent No. US 12,079,726 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following mapping below. Each corresponding limitation is either identical or does not have a patentable, nonobvious distinction unless otherwise noted. Instant Application 18/775,643 Patent No.: US 12,079,726 B2 Claim 21 Claims 8, 11, 13 Claim 22 Claim 8 Claim 23 Claim 8 Claim 24 Claim 10 Claim 25 Claim 12 Claim 26 Claim 13 Claim 27 Claim 9 Claim 28 Claims 15, 17 Claim 29 Claim 15 Claim 30 Claim 15 Claim 31 Claim 18 Claim 32 Claim 19 Claim 33 Claim 16 Claim 34 Claims 1, 4 Claim 35 Claim 1 Claim 36 Claim 1 Claim 37 Claim 3 Claim 38 Claim 5 Claim 39 Claim 6 Claim 40 Claim 2 Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-5, 9-13, 18 of U.S. Patent No. US 11,604,992 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following mapping below. Each corresponding limitation is either identical or does not have a patentable, nonobvious distinction unless otherwise noted. Instant Application 18/775,643 Patent No.: US 11,604,992 B2 Claim 21 Claims 1, 3 Claim 22 Claim 9 Claim 23 Claim 10 Claim 24 Claim 11 Claim 25 Claim 12 Claim 26 Claim 13 Claim 27 Claim 1 Claim 28 Claims 1, 2, 3 Claim 29 Claim 9 Claim 30 Claim 10 Claim 31 Claim 12 Claim 32 Claim 13 Claim 33 Claim 18 Claim 34 Claims 1, 3 Claim 35 Claim 1 Claim 36 Claim 1 Claim 37 Claim 2 Claim 38 Claim 4 Claim 39 Claim 5 Claim 40 Claim 1 Claim Rejections - 35 USC § 102 5. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 6. Claims 21-25, 27-38, and 40 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Zhou et al. (U.S. Patent Application Pub. No. US 20190354837 A1). Claim 21: Zhou teaches a method for generating a neural network architecture (i.e. the policy network uses (215) that network embedding to automatically generate adaptations to the neural network architecture configuration. In one or more embodiments, the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048]), the method comprising: sampling training data from a training data store (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]); determining, based on an initial probability distribution, a sample neural network architecture (i.e. the new layers may be generated in a stochastic way. FIG. 6 depicts a methodology to facilitate architecture configuration exploration using probability mass functions. Hence, in one or more embodiments, a goal of the insert LSTM is to define the probability mass function (p.m.f.) (e.g., PMFP 350-P) to sample (650) the features of the new layer to be generated. For each feature, mapping of the LSTM state output to the p.m.f. may be done by a lookup table 345. For example, if there are three (3) candidate values for the feature of convolution width, the LSTM state output determines three (3) probability values corresponding to them; para. [0062]) from a neural network architecture space (i.e. actions of scale and insert are mapped to a search space to define the neural network architectures … layer-by-layer search aims to find the optimal architecture with a search granularity of predefined layers, the neural network architecture is defined (705) by stacking these layers, an LSTM in the policy network chooses (710) the layer type and the corresponding hyperparameters (e.g., filter width); para. [0066, 0068]); performing an evaluation of the sampled training data using the sample neural network architecture (i.e. the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048, 0076, 0103]); updating, based on the evaluation, the initial probability distribution to generate an updated probability distribution (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0103]); determining whether termination criteria is satisfied (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077]); and when it is determined that the termination criteria is satisfied, generating a result neural network architecture based on the updated probability distribution (i.e. the updated policy network is used for the next episode … At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0077, 0103, 0105]). Claim 22: Zhou teaches the method of claim 21. Zhou further teaches wherein performing the evaluation of the sampled training data (i.e. the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048, 0076, 0103]) comprises: evaluating the sampled training data using the sampled neural network architecture to compute a gradient of a loss function associated with the sampled neural network architecture (i.e. The ADAM optimization algorithm was used for training each KWS model, a cross entropy (CE) loss function was used for training … The parameters θ of the policy network may be optimized via gradient descent; para. [0088, 0089, 0115]). Claim 23: Zhou teaches the method of claim 22. Zhou further teaches wherein updating the initial probability distribution to generate the updated probability distribution comprises: updating the initial probability distribution to generate the updated probability distribution based on the computed gradient of the loss function (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0077, 0103]). Claim 24: Zhou teaches the method of claim 21. Zhou further teaches wherein determining whether the termination criteria is satisfied comprises comparing a first accuracy of a neural network architecture associated with the initial probability distribution (i.e. Given a set of sample networks, performance curves for each network may be obtained. For each network xi, a validation accuracy ai and training time ti may be obtained; para. [0076, 0087, 0094]) and a second accuracy of a neural network architecture associated with the updated probability distribution (i.e. Each model was evaluated after training and an action is selected according to the current policy in order to transform the network. At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0062, 0077, 0103]) based on a predetermined threshold (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077, 0098]). Claim 25: Zhou teaches the method of claim 21. Zhou further teaches comprising: training the result neural network architecture (i.e. The top eight models from each episode were progressively selected as baseline models to the next episode. We train the best models for longer training time to get SOTA performance; para. [0096]) using training data from the training data store (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]). Claim 27: Zhou teaches the method of claim 21. Zhou further teaches comprising: when it is not determined that the termination criteria is satisfied, performing another iteration to further update the probability distribution (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0076, 0077]). Claim 28: Zhou teaches a method for generating a neural network architecture (i.e. the policy network uses (215) that network embedding to automatically generate adaptations to the neural network architecture configuration. In one or more embodiments, the adapted neural network is trained (220) to convergence, and the trained adapted neural network architecture may be evaluated (225) based upon one or more metrics (e.g., accuracy, memory footprint, power consumption, inference latency, etc.); para. [0048]), the method comprising: determining, based on an initial probability distribution, a sample neural network architecture (i.e. the new layers may be generated in a stochastic way. FIG. 6 depicts a methodology to facilitate architecture configuration exploration using probability mass functions. Hence, in one or more embodiments, a goal of the insert LSTM is to define the probability mass function (p.m.f.) (e.g., PMFP 350-P) to sample (650) the features of the new layer to be generated. For each feature, mapping of the LSTM state output to the p.m.f. may be done by a lookup table 345. For example, if there are three (3) candidate values for the feature of convolution width, the LSTM state output determines three (3) probability values corresponding to them; para. [0062]) from a neural network architecture space (i.e. actions of scale and insert are mapped to a search space to define the neural network architectures … layer-by-layer search aims to find the optimal architecture with a search granularity of predefined layers, the neural network architecture is defined (705) by stacking these layers, an LSTM in the policy network chooses (710) the layer type and the corresponding hyperparameters (e.g., filter width); para. [0066, 0068]); updating, based on an evaluation of sampled training data using the sample neural network architecture (i.e. The dataset split was also similar to [Zhang et al., 2017] that training, validation, and test sets have the ratio of 80:10:10 … An episode size of 5 and a batch size of 10 was used for all experiments, i.e. 10 child models are trained concurrently … Storage device(s) 1508 may also be used to store processed data or data to be processed in accordance with the disclosure; para. [0101, 0103, 0122]), the initial probability distribution to generate an updated probability distribution (i.e. the child networks are trained (1210) until convergence and a combination of performance and resource use are used (1215) as an immediate reward, as given in Eq. 3 (see also 1115 in FIG. 11). Rewards of a full episode (e.g., episode 1105 in FIG. 11) may be accumulated to train the policy network using the policy gradient to get an updated policy network (e.g., updated policy network 1120); para. [0076, 0103]); determining whether termination criteria is satisfied by comparing a first accuracy of a neural network architecture associated with the initial probability distribution (i.e. Given a set of sample networks, performance curves for each network may be obtained. For each network xi, a validation accuracy ai and training time ti may be obtained; para. [0076, 0087, 0094]) and a second accuracy of a neural network architecture associated with the updated probability distribution (i.e. Each model was evaluated after training and an action is selected according to the current policy in order to transform the network. At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0062, 0077, 0103]) based on a predetermined threshold (i.e. the updated policy network is used for the next episode. The number of episodes may be user-selected or may be based upon one or more stop conditions (e.g., runtime of RENA embodiment, number of iterations, convergence (or difference between iteration is not changing more than a threshold, divergence, and/or performance of the neural network meets criteria); para. [0077, 0098]); and when it is determined that the termination criteria is satisfied, generating a result neural network architecture based on the updated probability distribution (i.e. the updated policy network is used for the next episode … At the end of each episode, the policy was updated and the best 10 child models were used as the baseline for the new episode; para. [0077, 0103, 0105]). Claims 29-38 and 40 are similar in scope to Claims 21-25, 27 and are rejected under a similar rationale. Claim Rejections – 35 USC § 103 7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 8. Claims 26 and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Vasudevan et al. (U.S. Patent Application Pub. No. US 20190026639 A1). Claim 26: Zhou teaches the method of claim 21. Zhou further teaches wherein the initial probability distribution is determined based on a surrogate neural network architecture (i.e. a search can start with any baseline models, a well-designed or even a rudimentary one … a policy network embodiment 300 uses a network embedding 320 to represent the input neural network configuration 302; para. [0058, 0059, 0061, 0062]). Zhou does not explicitly teach was determined using a dataset surrogate, wherein the dataset surrogate comprises a different set of training data from the training data store than a set of training data used when sampling training data from a training data store. However, Vasudevan teaches a surrogate neural network architecture that was determined using a dataset surrogate (i.e. the system can effectively determine the architecture of the convolutional cells on a smaller data set and then re-use the same cell architecture across a range of data and computational scales … The neural architecture search system 100 is a system that obtains training data 102 for training a convolutional neural network to perform a particular task and a validation set 104 for evaluating the performance of the convolutional neural network on the particular task and uses the training data 102 and the validation set 104 to determine an network architecture for a child CNN that is configured to perform the image processing task … Once trained values of the controller parameters have been determined, i.e., once the training of the controller neural network 110 has satisfied some termination criteria, the system determines a final architecture for the first convolutional cell (and any other convolutional cells that are defined by the output sequences generated by the controller neural network); para. [0006, 0016, 0035]), wherein the dataset surrogate comprises a different set of training data from the training data store than a set of training data used when sampling training data from a training data store (i.e. the system can effectively determine the architecture of the convolutional cells on a smaller data set and then re-use the same cell architecture across a range of data and computational scales. In particular, the system can effectively employ the resulting learned architecture to perform image processing tasks with reduced computational budgets that match or outperform streamlined architectures targeted to mobile and embedded platforms; para. [0006, 0038]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Zhou to include the feature of Vasudevan. One would have been motivated to make this modification because the architecture learned using a smaller dataset can be reused across different data scales, thereby reducing the computational resources. Claim 39 is similar in scope to Claim 26 and is rejected under a similar rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Sun et al. (Pub. No. US 20200082275 A1), Disclosed are a neural network architecture search apparatus and method and a computer readable recording medium. The neural network architecture search method comprises: defining a search space used as a set of architecture parameters describing the neural network architecture; performing sampling on the architecture parameters in the search space based on parameters of a control unit, to generate at least one sub-neural network architecture; performing training on each sub-neural network architecture by minimizing a loss function including an inter-class loss and a center loss. It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TAN H TRAN/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Jul 17, 2024
Application Filed
Oct 08, 2025
Response after Non-Final Action
Sep 09, 2026
Non-Final Rejection mailed — §102, §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748960
Analog Hardware Realization of Neural Networks
5y 7m to grant Granted Sep 29, 2026
Patent 12718079
Systems and Methods for Generating Libraries for Hardware Realization of Neural Networks
5y 5m to grant Granted Aug 25, 2026
Patent 12718088
DESIGNING LADDER AND LAGUERRE ORTHOGONAL RECURRENT NEURAL NETWORK ARCHITECTURES INSPIRED BY DISCRETE-TIME DYNAMICAL SYSTEMS
4y 8m to grant Granted Aug 25, 2026
Patent 12688413
METHODS FOR RELIABLE OVER-THE-AIR COMPUTATION WITH PULSES FOR DISTRIBUTED LEARNING AND WITH FEDERATED EDGE LEARNING WITHOUT CHANNEL STATE INFORMATION
4y 1m to grant Granted Jul 21, 2026
Patent 12682274
MODEL INTEGRATION APPARATUS, MODEL INTEGRATION METHOD, COMPUTER-READABLE STORAGE MEDIUM STORING A MODEL INTEGRATION PROGRAM, INFERENCE SYSTEM, INSPECTION SYSTEM, AND CONTROL SYSTEM
5y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
94%
With Interview (+32.6%)
3y 6m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 320 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month