Prosecution Insights
Last updated: October 02, 2026
Application No. 18/214,188

GLOBAL OPTIMIZATION FOR NEURAL NETWORK TRAINING

Final Rejection §103
Filed
Jun 26, 2023
Examiner
BREENE, PAUL J
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
2 (Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
11m
Est. Remaining
77%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
42 granted / 68 resolved
+6.8% vs TC avg
Strong +15% interview lift
Without
With
+15.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
14 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
27.7%
-12.3% vs TC avg
§103
47.7%
+7.7% vs TC avg
§102
8.4%
-31.6% vs TC avg
§112
15.4%
-24.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 68 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over US Pre-Grant Patent 2020/0125930 (Martin et al; Martin) in view of “Training multilayer neural networks using fast global learning algorithm- least-squares and penalized optimization methods,” Neurocomputing 25 (1999): 115-131, Cho et al; Cho. Regarding claim 1, and analogous claims 10 and 18: Martin teaches: 1. receiving a data set for training a machine learning model to perform a recognition task; (Martin, ¶0051) “A new task was created by randomly generating a permutation mask and applying it to each of the digits in the dataset [i.e. receiving a data set for training a machine learning model].” (Martin, ¶0051) “The artificial neural networks of the present disclosure, and the methods of retraining the artificial neural networks according to various embodiments of the present disclosure, were also tested with a variant of the MNIST optical character recognition problem [i.e. to perform a recognition task;].” 2. training the machine learning model, by performing an optimization that accelerates the training, wherein the optimization comprises at least: (Martin, ¶0051) “The parameter λ in Equation 2 above was set to 10.0. It was determined that smaller values of λ resulted in the complexity term dominating the fitness, which resulted in a fairly simple fitness landscape with the global optimum being achieved by adding only 1 to 3 neurons at any layer [i.e. training the machine learning model, by performing an optimization that accelerates the training, wherein the optimization comprises at least:].” 3. searching for a minimum value of a loss function; (Martin, ¶0051) “Setting λ=10.0 provided a better balance between accuracy and complexity, and consequently, a more challenging optimization problem with many good, but suboptimal, local minima [i.e. searching for a minimum value of a loss function;].” 4. and updating the machine learning model with parameters identified at the global minimum. (Martin, ¶0051) “In this setting, the global optimum is achieved by adding 17 new neurons to the first hidden layer and no new neurons to the second and third hidden layers. However, good, but suboptimal, local minima can be achieved by adding new neurons to only the second or third hidden layers [i.e. and updating the machine learning model with parameters identified at the global minimum].” Martin does not explicitly teach: 1. responsive to finding a local minimum, adding an additional term to the loss function and continuing to find another local minimum until a criterion is met, the additional term added to the loss function changing a loss surface of the loss function; 2. and identifying a global minimum having a lowest minimum value among the found local minima; Cho teaches: 1. responsive to finding a local minimum, adding an additional term to the loss function and continuing to find another local minimum until a criterion is met, the additional term added to the loss function changing a loss surface of the loss function; (Cho, pg. 119, Sect. 3.1, Eq. 11) “As our aforementioned statements, the minimization by the gradient descent optimization suffers from the problem of the local minima. Therefore, it is advisable to modify the new cost function, E(W) by including a penalty function to provide a search out of the local minima when the convergence gets stuck [i.e. responsive to finding a local minimum, adding an additional term to the loss function].” (Cho, pg. 115, Abstract) “The penalty term superimposes into the error surface, which likely to provide a way of escape from the local minima when the convergence stalls. The choice and adjustment for the penalty factor are also derived to demonstrate the effect of the penalty term and to ensure the convergence of the algorithm [i.e. the additional term added to the loss function changing a loss surface of the loss function;].” 2. and identifying a global minimum having a lowest minimum value among the found local minima; (Cho, pg. 118, Sect. 2, Eq. 1) “The objective is to find the global minimum solution, wGM which minimizes E(w)) where I is the domain of the state variables over which one seeks the global minimum and I is assumed to be compact and connected [i.e. and identifying a global minimum having a lowest minimum value among the found local minima;].” One of ordinary skill in the art at the time the invention was filed would have been motivated to modify Martin’s machine-learning training optimization to employ Cho’s penalty-based optimization technique when convergence becomes trapped at a local minimum, because Cho expressly recognizes that gradient-descent optimization suffers from convergence at local minima and teaches modifying the cost function by adding a penalty function to provide a search out of the local minimum. A skilled artisan would therefore have been motivated to make the modification to improve the ability of Martin’s optimization process to escape undesirable local minima and continue searching toward the global minimum, thereby improving convergence performance, as Cho explains that its proposed algorithm provides “the fastest rate of convergence” compared with the conventional algorithms tested and is capable of escaping local minima (Cho, pg. 124). Regarding claim 2, and analogous claims 11 and 19: The combination of Martin and Cho teach the method of claim 1. Martin teaches: 1. wherein the machine learning model includes a deep neural network and the optimization includes a descent-based optimization. (Martin, ¶0011) “In general, the artificial neural network can have any suitable number of hidden layers, and the methods of the present disclosure can add any suitable number of nodes to any of the hidden layers [i.e. wherein the machine learning model includes a deep neural network].” (Martin, ¶0019) “Training the artificial neural network on the data from the new task may include minimizing a loss function with stochastic gradient descent [i.e. and the optimization includes a descent-based optimization].” One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Regarding claim 3, and analogous claims 12 and 20: The combination of Martin and Cho teach the method of claim 1. Martin teaches: 1. wherein the additional term is a Gaussian bias centered around the local minimum. (Martin, ¶0039) “In one or more embodiments, the task 120 utilizes samples only from the probability distributions P.sub.Z1(Z1|X1) and P.sub.Z2(Z2|X1), and therefore the task 120 does not require closed-form expressions for the probability distributions, which may be, or may approximately be, Gaussian functions [i.e. wherein the additional term is a Gaussian bias centered around the local minimum].” One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Regarding claim 4, and analogous claim 13: The combination of Martin and Cho teach the method of claim 1. Cho teaches: 1. wherein the criterion includes a threshold number of local minima. (Cho, pgs. 121-122, Sect. 3.3) PNG media_image1.png 327 518 media_image1.png Greyscale Examiner notes that steps 5–7 repeatedly evaluate the optimization state and return to Step 5 until the stopping criterion terminates training. Cho further expressly contemplates an optimization landscape containing a finite plurality of minima, stating that “the problem has three minima, one of which is the global minimum,” and searches through the minima to identify the global minimum. Thus, under the broadest reasonable interpretation of “threshold number,” Cho’s iterative search through a finite plurality of local minima subject to a stopping criterion teaches or at least suggests a criterion encompassing a threshold number of local minima. One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Regarding claim 5, and analogous claim 14: The combination of Martin and Cho teach the method of claim 1. Cho teaches: 1. wherein the additional term is added to the loss function until the local minimum is filled according to a threshold wherein the local minimum is no longer recognized as the minimum value of the loss function. (Cho, pg. 121, Sect. 3.2) “When the convergence is stuck in local minimum, the penalized optimization is introduced. μ(t) decreases gradually by 0.3% of the μ(t-1) to provide an uphill force to escape from the local minimum [i.e. wherein the additional term is added to the loss function until the local minimum is filled]. After running the penalized optimization for few iterations, the μ(t) should be increased by 1% of the μ(t-1) to diminish the effect of the penalty term when the training error starts to decrease [i.e. according to a threshold wherein the local minimum is no longer recognized as the minimum value of the loss function]. One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Regarding claim 6, and analogous claim 15: The combination of Martin and Cho teach the method of claim 1. Martin teaches: 1. wherein the method further includes storing the additional term. (Martin, ¶0045) “In one or more embodiments, the task 160 of training the artificial neural network 200 includes minimizing the following loss function using stochastic gradient descent [i.e. wherein the additional term is added to the loss function until the local minimum is filled].” Examiner notes that in order for the training to function properly, the additional term would have to be stored. One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Chiang. The motivation is the same as claim 1. Regarding claim 7: The combination of Martin and Cho teach the method of claim 1. Martin teaches: 1. further comprising reconstructing an original landscape of the loss function by accessing all stored additional terms added to the loss function and removing the accessed additional terms from the loss function (Martin, ¶0045) “The third term of the loss function (Equation 2) not only helps prevent catastrophic forgetting of old tasks, but also enables some drift in the hidden distributions, which promotes integration of information from old and new tasks, thus reducing the required size of the artificial neural network 200 (i.e., minimizing or at least reducing the number of nodes and connections) for a given performance level [i.e. further comprising reconstructing an original landscape of the loss function by accessing all stored additional terms added to the loss function and removing the accessed additional terms from the loss function].” Examiner interprets the integration of information from old or new tasks as removing and adding terms. Further, integration of information from old or new tasks would constitute “all” stored additional terms. One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Regarding claim 9, and analogous claim 17: The combination of Martin and Cho teach the method of claim 1. Martin teaches: 1. wherein the computer is further caused to use the updated machine learning model in performing a recognition task. (Martin, ¶0051) “A new task was created by randomly generating a permutation mask and applying it to each of the digits in the dataset. The permutation mask was created by randomly selecting two non-intersecting sets of pixel indices, and then swapping the corresponding pixels in each image [i.e. wherein the computer is further caused to use the updated machine learning model in performing a recognition task].” One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin in view of Cho. The motivation is the same as claim 1. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over US Pre-Grant Patent 2020/0125930 (Martin et al; Martin) in view of “Training multilayer neural networks using fast global learning algorithm- least-squares and penalized optimization methods,” Neurocomputing 25 (1999): 115-131, Cho et al; Cho, further in view of US Pre-Grant Patent 2011/0137830 (Ozyurt et al; Ozyurt). Regarding claim 8, and analogous claim 16: The combination of Martin and Cho teach the method of claim 1. Cho teaches: 1. [wherein multiple instances of the optimization are performed in parallel at different initialization points of] a loss surface of the loss function. (Cho, pg. 115, Abstract) “The penalty term superimposes into the error surface, which likely to provide a way of escape from the local minima when the convergence stalls. The choice and adjustment for the penalty factor are also derived to demonstrate the effect of the penalty term and to ensure the convergence of the algorithm [i.e. a loss surface of the loss function;].” Neither Martin and Cho teach: 1. wherein multiple instances of the optimization are performed in parallel at different initialization points [of a loss surface of the loss function.] (Ozyurt, ¶0059) “Note that the point set may include one or more start points that may be used by the one or more sub-solvers [i.e. wherein multiple instances of the optimization are performed in parallel at different initialization points]” One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Martin and Cho in view of Ozyurt. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL JUSTIN BREENE whose telephone number is (571)272-6320. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web- based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on 303-297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786 9199 (IN USA OR CANADA) or 571-272-1000. /P.J.B./ Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Jun 26, 2023
Application Filed
Mar 12, 2026
Non-Final Rejection mailed — §103
May 20, 2026
Interview Requested
Jun 09, 2026
Response Filed
Aug 27, 2026
Final Rejection mailed — §103
Sep 14, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748989
DYNAMIC AUGMENTATION BASED ON DATA SAMPLE HARDNESS
6y 2m to grant Granted Sep 29, 2026
Patent 12743641
ABNORMALITY DETECTION BASED ON CAUSAL GRAPHS REPRESENTING CAUSAL RELATIONSHIPS OF ABNORMALITIES
5y 7m to grant Granted Sep 22, 2026
Patent 12737586
Methods and apparatuses for compressing parameters of neural networks
1y 4m to grant Granted Sep 15, 2026
Patent 12726184
Accelerated Learning In Neural Networks Incorporating Quantum Unitary Noise And Quantum Stochastic Rounding Using Silicon Based Quantum Dot Arrays
4y 9m to grant Granted Sep 01, 2026
Patent 12725008
TECHNIQUES FOR INPUT CLASSIFICATION AND RESPONSE USING GENERATIVE NEURAL NETWORKS
4y 5m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
77%
With Interview (+15.4%)
4y 2m (~11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 68 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month