Prosecution Insights
Last updated: October 01, 2026
Application No. 18/530,683

METHOD AND APPARATUS WITH DISTRIBUTED TRAINING OF NEURAL NETWORK

Final Rejection §103
Filed
Dec 06, 2023
Priority
Jul 20, 2023 — RE 10-2023-0094737
Examiner
NILSSON, ERIC
Art Unit
Tech Center
Assignee
Seoul National University R&DB Foundation
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
430 granted / 520 resolved
+22.7% vs TC avg
Strong +18% interview lift
Without
With
+17.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
23 currently pending
Career history
535
Total Applications
across all art units

Statute-Specific Performance

§101
27.2%
-12.8% vs TC avg
§103
41.8%
+1.8% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
8.5%
-31.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 520 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to claims filed 21 July 2026 for application 18530683 filed 06 December 2023. Currently claims 1-20 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1, 5, 6, 10-13, 16, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Volodarskiy et al. (US 20200175354 A1) in view of Yi et al. (Fast Training of Deep Learning Models over Multiple GPUs). Regarding claims 1, 12, 13, and 19, Volodarskiy teaches: A processor-implemented method, the method comprising: while training a neural network (NN) using a current training mode selected from a plurality of training modes for training of the NN, measuring time data of a plurality of sub-operations for the training of the NN (“An underappreciated aspect of AutoML is that trying all possible algorithms and their hyperparameter values is an infinite time problem. Because of this infinite time problem, all AutoML tools and well-known academic algorithms constrain the time for algorithm selection and training. However, these AutoML tools do not formulate the problem as selecting the best algorithm based on the time constraint. Thus, in an embodiment, a method is disclosed that comprises using at least one hardware processor to select a plurality of machine-learning algorithms, generate a batch of trials from the plurality of machine-learning algorithms, begin executing at least a portion of the batch of trials, and, during execution of the batch of trials, provide intermediate evaluation results. It may involve estimating a training time and accuracy for two or more models represented in the batch of trials, and then selecting the best algorithm and hyperparameter settings to train from an available set of algorithm/hyperparameter setting combinations based on a time constraint set by the user. The method may use both an estimate of the training time for each algorithm and an estimated accuracy of the algorithm and settings to make the selection. For example, the method may use a combination of the estimated training time and the a priori (pre-training) estimated accuracy of the algorithm to choose the next best algorithm to train.” [0049], claim 6 model may be neural network); based on the time data, determining a computation time to perform computation operations among the plurality of sub-operations and a communication time to perform communication operations among the plurality of sub-operations [0049]; based on a comparison result of the computation time and the communication time, selecting a next training mode from the plurality of training modes (“For example, the method may use a combination of the estimated training time and the a priori (pre-training) estimated accuracy of the algorithm to choose the next best algorithm to train.” [0049]); and training the NN based on the next training mode (“For example, the method may use a combination of the estimated training time and the a priori (pre-training) estimated accuracy of the algorithm to choose the next best algorithm to train.” [0049]). Claims 13 and 19 also recite: processing modules configured to execute workloads corresponding to the plurality of sub-operations. Volodarskiy discloses: processing modules configured to execute workloads corresponding to the plurality of sub-operations (“The infrastructure may comprise a platform 110 (e.g., one or more servers) which hosts and/or executes one or more of the various functions, processes, methods, and/or software modules described herein. Platform 110 may comprise dedicated servers, or may instead comprise cloud instances, which utilize shared resources of one or more servers. These servers or cloud instances may be collocated and/or geographically distributed. Platform 110 may also comprise or be communicatively connected to a server application 112 and/or one or more databases 114. In addition, platform 110 may be communicatively connected to one or more user systems 130 via one or more networks 120. Platform 110 may also be communicatively connected to one or more external systems 140 (e.g., other platforms, websites, etc.) via one or more networks 120.” [0019]). Volodarksiy does not explicitly disclose: corresponding to respective update ranges of different portion sizes of the NN. Yi teaches: corresponding to respective update ranges of different portion sizes of the NN. (“To build the communication cost model, we gather tensors across the same source-destination device pairs into one group. For each group, we use linear regression to obtain a linear model: tensor size vs. transfer time. In each update of the cost model, newly collected data are fed and parameters of the linear model are re-computed. The models capture available bandwidth and potential congestion along each device-device path. Strategy Calculator. It is the key component to carry out the algorithms that we will discuss in Sec. 5. During the pre training stage, it calculates device placement, execution order and operation partition lists, and obtains the cost models. During the normal training stage, it periodically activates the profiler, updates the cost models, and recalculates new strategies. If the estimated per-iteration training time with the new strategies (among output of our DPOS algorithm) is smaller than that of previous strategies, the new strategies are activated.” P109 ¶1-2) Volodarskiy and Yi are in the same field of endeavor of training distributed neural networks and are analogous. Volodarskiy discloses a method for determining training time and changing training methods depending on user constraints. Yi discloses changing training modes depending on update sizes and ranges. It would have been obvious to modify the training methods as taught by Volodarskiy to utilize the adaptive training methods as taught by Yi to yield the predictable results of reduced training time. Regarding claims 5 and 16, Volodarskiy teaches: The method of claim 1, wherein the plurality of sub-operations comprises any one or any combination of any two or more of a backward computation operation related to backward propagation, a gradient communication operation related to sharing of a layer gradient, an update computation operation related to model update, and a parameter communication operation related to sharing of a model parameter (“It may involve estimating a training time and accuracy for two or more models represented in the batch of trials, and then selecting the best algorithm and hyperparameter settings to train from an available set of algorithm/hyperparameter setting combinations based on a time constraint set by the user.” [0049]). Regarding claim 6, Volodarskiy teaches: The method of claim 5, wherein the determining of the computation time and the communication time comprises, based on the time data, recording first temporary data of any one or any combination of any two or more of the backward computation operation, the gradient communication operation, the update computation operation, and the parameter communication operation in a timetable for each layer of the NN (“It may involve estimating a training time and accuracy for two or more models represented in the batch of trials, and then selecting the best algorithm and hyperparameter settings to train from an available set of algorithm/hyperparameter setting combinations based on a time constraint set by the user.” [0049]). Regarding claims 10, 18 and 20, Volodarskiy teaches: The method of claim 1, wherein the selecting of the next training mode comprises: in response to a value of the computation time being larger among the computation time and the communication time, selecting the next training mode such that the value of the computation time decreases; and in response to the value of the computation time being larger among the computation time and the communication time, selecting the next training mode such that the value of the computation time increases (“It may involve estimating a training time and accuracy for two or more models represented in the batch of trials, and then selecting the best algorithm and hyperparameter settings to train from an available set of algorithm/hyperparameter setting combinations based on a time constraint set by the user.” [0049]). Regarding claim 11, Volodarskiy teaches: The method of claim 1, further comprising: predicting a change in total training time according to the next training mode based on dependency between computation operations and communication operations of a plurality of layers of the NN; and in response to the total training time increasing according to the next training mode, selecting an alternative training mode of the next training mode from the plurality of training modes (“It may involve estimating a training time and accuracy for two or more models represented in the batch of trials, and then selecting the best algorithm and hyperparameter settings to train from an available set of algorithm/hyperparameter setting combinations based on a time constraint set by the user.” [0049]). Claim(s) 2-3 and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Volodarskiy in view of Yi and further in view of Mopur et al. (US 20200151619 A1). Regarding claims 2 and 14, Volodarskiy does not explicitly disclose, however, Mopur teaches: The method of claim 1, wherein the plurality of training modes are distinguished from each other according to an update range of the NN by each processing module used for the training of the NN [0045-46] either a full range historical and new retraining or only new retraining can be used). Volodarskiy, Yi and Mopur are in the same field of endeavor of training ML models and are analogous. Volodarskiy discloses a method for determining training time and changing training methods depending on user constraints. Yi discloses changing training modes depending on update sizes and ranges. Mopur teaches training using historical and new or only new data depending on metrics. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the training modification due to time constraints as disclosed by Volodarskyi and Yi with the know training using various amounts and types of data as disclosed by Mopur to yield predictable results of faster and better training. Regarding claims 3 and 15, Volodarskiy does not explicitly disclose, however, Mopur teaches: The method of claim 1, wherein the plurality of training modes comprises any one or any combination of any two or more of: a first training mode in which full update of a corresponding model of the NN is performed by each of processing modules used for the training of the NN (“Another form of training the impact of data drift may indicate the need to add additional prediction metrics to the original training regime. For example, in an image classification scenario, a classification done by the production model version (i.e., v1.A) may be with low confidence and be erroneous. The result could be manually inspected by a user (i.e., data expert) and manually reclassifies the image. This new classification data is therein used to train the model. This type of user controlled training would apply additional prediction metrics to the original production ML model 208 in an attempt to identify the unknown relationships impacting performance In various embodiments, more than one user controlled training can be launched.” [0046]); a second training mode in which partial update of 1/N of the corresponding model is performed by each of the processing modules (“Other impacts may indicate the need to perform new data training. For example, in some embodiments a sequence of negative or low correlation values may indicate that the drift has resulted in a drastic change in performance. In such cases, utilizing only new streaming data may be preferable, as the historical data has consistently resulted in poor predicted performance. Accordingly, training may be performed on the first production ML model 208 as deployed may be trained only with new streaming data to try and achieve a faster improvement in performance. This would result in a new version of the production ML model (production ML model 208.2).” [0045]); and a third training mode in which partial update of 1/M of the corresponding model is performed by each of the processing modules, and wherein the N represents a total number of the processing modules, and the M represents an integer greater than 1 and smaller than the N ([0045-46]. Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Allowable Subject Matter Claims 4, 7-9 and 17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. None of the prior art of record teaches or discloses modifying the training mode and/or training data on a per layer basis in a neural network. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC NILSSON whose telephone number is (571)272-5246. The examiner can normally be reached M-F: 7-3. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571)-272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ERIC NILSSON/ Primary Examiner, Art Unit 2151
Read full office action

Prosecution Timeline

Dec 06, 2023
Application Filed
Jun 03, 2026
Non-Final Rejection mailed — §103
Jul 21, 2026
Interview Requested
Jul 21, 2026
Response Filed
Aug 11, 2026
Applicant Interview (Telephonic)
Aug 11, 2026
Examiner Interview Summary
Aug 25, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749019
ACCELERATED LEARNING FROM SPATIO-TEMPORAL DATA
3y 2m to grant Granted Sep 29, 2026
Patent 12743604
FUNCTION-BASED ACTIVATION OF MEMORY TIERS
4y 0m to grant Granted Sep 22, 2026
Patent 12737638
TRANSFER LEARNING OF MACHINE LEARNING MODEL IN DISTRIBUTED NETWORK
3y 9m to grant Granted Sep 15, 2026
Patent 12737684
MODEL-SPECIFIC SYNTHETIC DATA GENERATION FOR MACHINE LEARNING MODEL TRAINING
3y 3m to grant Granted Sep 15, 2026
Patent 12737628
INTELLIGENT RECOGNITION AND ALERT METHODS AND SYSTEMS
3y 2m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
99%
With Interview (+17.6%)
3y 1m (~3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 520 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month