Prosecution Insights
Last updated: October 04, 2026
Application No. 18/322,373

NEURAL NETWORK TRAINING METHOD AND APPARATUS

Final Rejection §102
Filed
May 23, 2023
Priority
Nov 23, 2020 — CN 202011322834.6 +1 more
Examiner
LANE, THOMAS BERNARD
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
Huawei Technologies Co., Ltd.
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
13 granted / 17 resolved
+21.5% vs TC avg
Minimal +5% lift
Without
With
+5.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
20 currently pending
Career history
34
Total Applications
across all art units

Statute-Specific Performance

§101
22.4%
-17.6% vs TC avg
§103
45.9%
+5.9% vs TC avg
§102
15.3%
-24.7% vs TC avg
§112
12.9%
-27.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 17 resolved cases

Office Action

§102
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. CN202010745395.3, filed on 11/23/2020. Information Disclosure Statement The information disclosure statement (IDS) submitted on 10/04/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments Applicant's arguments filed 06/09/2026 regarding the rejection under 35 USC 102 have been fully considered but they are not persuasive. Applicant argues, see especially page 9-10, that claims 1, 8, and 15, are patent eligible because “The Office Action rejected Claim 1 as being anticipated by Xiao. Claim 1 requires, in part, "obtaining sampling probability distribution wherein the sampling probability distribution represents a probability that each of the M groups of parameters is sampled in each training iteration step." At least these features of claim 1 are not taught, disclosed, or suggested by the cited prior art. The Office Action points to Xiao's "freezing rate" (e.g., page 1229-1230) as teaching this sampling probability distribution. However, Xiao's freezing rate is a calculated metric based on whether gradients are "likely to be canceled out" in order to determine if a layer should be frozen. This is fundamentally different from a sampling probability distribution that represents a probability that a group of parameters is sampled. In Claim 1, the sampling probability distribution is obtained and used to sample a parameter group, which is then frozen or stopped from updating. This implies a stochastic or distribution-based selection process for the purpose of freezing. Xiao, by contrast, uses a deterministic heuristic based on training performance (gradient cancellation) to identify layers for freezing. Xiao does not teach a distribution where each of M groups has a specific probability of being "sampled" in the sense of the claim. Because Xiao fails to teach the "sampling probability distribution" used to sample groups for freezing as required by Claim 1, Xiao does not anticipate Claim 1. Independent Claims 8 and 15, which recite similar limitations in the context of an apparatus and a computer- readable medium, are also patentable over Xiao for the same reasons. Dependent claims 2-7, 9-14, and 16-20 are patentable by virtue of their dependence on the independent claims. Reconsideration of the ground of rejection and indication of the allowability of all pending claims are respectfully requested.” Examiner respectfully disagrees. Xiao teaches the use of the freezing rate to determine the probability of a layer being sampled by multiple tasks that would cause the weights to move towards 0 or cancel out the learning, this represents a “sampling probability distribution represents a probability that each of the M groups of parameters is sampled in each training iteration step”. The Xiao reference further teaches a plurality of these freezing rates are taken for all the weights and layers which are parameters and groups of parameters. Under the broadest reasonable interpretation of the claims as they are presently stated Xiao teaches these limitations. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xiao et al. “Fast Deep Learning Training Through Intelligently Freezing Layers”, 06/17/2019. Regarding Claim 1 Xiao teaches A neural network training method applied to a neural network training apparatus, the method comprising: obtaining a to-be-trained neural network; (Xiao, page 1230-1231, section IV, teaches the obtaining of different models to be trained by the neural network training apparatus.) grouping parameters of the to-be-trained neural network, to obtain M groups of parameters, wherein M is a positive integer greater than or equal to 1; (Xiao, page 1228, section III – A and C, teaches the configuring and use of neural network layers which are groups of weights and parameters used by the neural network to generate its decisions.) obtaining sampling probability distribution and training iteration step arrangement, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e. groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution). Further Xiao, page 1229 -1230, section III – A, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement).) wherein the sampling probability distribution represents a probability that each of the M groups of parameters is sampled in each training iteration step, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen) and the training iteration step arrangement comprises interval arrangement and periodic arrangement; (Xiao, page 1229 -1230, section III – A, and B, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement). This includes the determining of which layers should be frozen at each epoch.) freezing or stopping updating a sampled parameter group based on the sampling probability distribution and the training iteration step arrangement; (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen and training the to-be-trained neural network based on the parameter group that is frozen or stopped updating. (Xiao, page 1229 -1230, section III – B, C, and E, teach the freezing of layers (i.e. parameter groups) in a neural network and then continuing training the neural network with those frozen layers) Regarding Claim 2 Xiao teaches The method according to claim 1, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and freezing the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation and parameter update are not performed on the mth group of parameters to the first group of parameters. ; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen. When a layer is chosen to be frozen it is frozen for subsequent epochs and the gradient calculation and parameter updates are not performed on that layer.) Regarding Claim 3 Xiao teaches The method according to claim 1, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and stopping updating the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation is performed and parameter update are not performed on the mth group of parameters to the first group of parameters. (Xiao, page 1229 -1230, section III – B, C, E, and D and Algorithm 1, Fig, $, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be Stopped updating. In figure 4 it is shown that when a layer is frozen if the layer after it is unfrozen the gradient for the frozen layer will still be calculated but the parameters will not be updated (i.e. stopping updating).) Regarding Claim 4 Xiao teaches The method according to claim 2, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 5 Xiao teaches The method according to claim 2, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 6 Xiao teaches The method according to claim 2, wherein in response to the training iteration step arrangement being the periodic arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining that a quantity of first iteration steps is M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. iteration step arrangement))) and determining a first period based on the quantity of first iteration steps and a first proportion, wherein the first period comprises the first iteration step and an iteration step to be trained on the entire network, the first proportion is a proportion of the first iteration step in the first period, and the first iteration step is last (M-1) iteration steps in the first period. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the periodic freezing of layers based on how many epochs have been performed (i.e. iteration step) as the epochs are performed the freezing rate increases for each individual layer and when the set period of epochs is hit the layers indicated by the freezing rate will be frozen or stopped till the ending of the training epochs (i.e. last iteration step in the first period.) ) Regarding Claim 7 Xiao teaches The method according to claim 2, wherein in response to the training iteration step arrangement being the periodic arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining that a quantity of first iteration steps is M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. iteration step arrangement))) and determining a first period based on the quantity of first iteration steps and a first proportion, wherein the first period comprises the first iteration step and an iteration step to be trained on the entire network, the first proportion is a proportion of the first iteration step in the first period, and the first iteration step is last (M-1) iteration steps in the first period. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the periodic freezing of layers based on how many epochs have been performed (i.e. iteration step) as the epochs are performed the freezing rate increases for each individual layer and when the set period of epochs is hit the layers indicated by the freezing rate will be frozen or stopped till the ending of the training epochs (i.e. last iteration step in the first period.) ) Regarding Claim 8 Xiao teaches A neural network training apparatus, comprising: a memory having computer-executable instructions stored thereon; and a processor configured to execute the computer-executable instructions in the memory to facilitate the following being performed by the apparatus: obtaining a to-be-trained neural network; (Xiao, page 1230-1231, section IV, teaches the obtaining of different models to be trained by the neural network training apparatus.) grouping parameters of the to-be-trained neural network, to obtain M groups of parameters, wherein M is a positive integer greater than or equal to 1; (Xiao, page 1228, section III – A and C, teaches the configuring and use of neural network layers which are groups of weights and parameters used by the neural network to generate its decisions.) obtaining sampling probability distribution and training iteration step arrangement, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution). Further Xiao, page 1229 -1230, section III – A, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement).) wherein the sampling probability distribution represents a probability that each of the M groups of parameters is sampled in each training iteration step, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen) and the training iteration step arrangement comprises interval arrangement and periodic arrangement; (Xiao, page 1229 -1230, section III – A, and B, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement). This includes the determining of which layers should be frozen at each epoch.) freezing or stopping updating a sampled parameter group based on the sampling probability distribution and the training iteration step arrangement; (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen and training the to-be-trained neural network based on the parameter group that is frozen or stopped updating. (Xiao, page 1229 -1230, section III – B, C, and E, teach the freezing of layers (i.e. parameter groups) in a neural network and then continuing training the neural network with those frozen layers) Regarding Claim 9 Xiao teaches The apparatus according to claim 8, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and freezing the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation and parameter update are not performed on the mth group of parameters to the first group of parameters. ; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen. When a layer is chosen to be frozen it is frozen for subsequent epochs and the gradient calculation and parameter updates are not performed on that layer.) Regarding Claim 10 Xiao teaches The apparatus according to claim 8, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and stopping updating the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation is performed and parameter update are not performed on the mth group of parameters to the first group of parameters. (Xiao, page 1229 -1230, section III – B, C, E, and D and Algorithm 1, Fig, $, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be Stopped updating. In figure 4 it is shown that when a layer is frozen if the layer after it is unfrozen the gradient for the frozen layer will still be calculated but the parameters will not be updated (i.e. stopping updating).) Regarding Claim 11 Xiao teaches The apparatus according to claim 9, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 12 Xiao teaches The apparatus according to claim 10, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 13 Xiao teaches The apparatus according to claim 9, wherein in response to the training iteration step arrangement being the periodic arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining that a quantity of first iteration steps is M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. iteration step arrangement))) and determining a first period based on the quantity of first iteration steps and a first proportion, wherein the first period comprises the first iteration step and an iteration step to be trained on the entire network, the first proportion is a proportion of the first iteration step in the first period, and the first iteration step is last (M-1) iteration steps in the first period. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the periodic freezing of layers based on how many epochs have been performed (i.e. iteration step) as the epochs are performed the freezing rate increases for each individual layer and when the set period of epochs is hit the layers indicated by the freezing rate will be frozen or stopped till the ending of the training epochs (i.e. last iteration step in the first period.) ) Regarding Claim 14 Xiao teaches The apparatus according to claim 10, wherein in response to the training iteration step arrangement being the periodic arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining that a quantity of first iteration steps is M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. iteration step arrangement))) and determining a first period based on the quantity of first iteration steps and a first proportion, wherein the first period comprises the first iteration step and an iteration step to be trained on the entire network, the first proportion is a proportion of the first iteration step in the first period, and the first iteration step is last (M-1) iteration steps in the first period. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the periodic freezing of layers based on how many epochs have been performed (i.e. iteration step) as the epochs are performed the freezing rate increases for each individual layer and when the set period of epochs is hit the layers indicated by the freezing rate will be frozen or stopped till the ending of the training epochs (i.e. last iteration step in the first period.) ) Regarding Claim 15 Xiao teaches A non-transitory computer-readable storage medium, wherein the computer-readable medium stores program code executed by a device, and the program code, upon execution the device, facilitating performance of the following: obtaining a to-be-trained neural network; (Xiao, page 1230-1231, section IV, teaches the obtaining of different models to be trained by the neural network training apparatus.) grouping parameters of the to-be-trained neural network, to obtain M groups of parameters, wherein M is a positive integer greater than or equal to 1; (Xiao, page 1228, section III – A and C, teaches the configuring and use of neural network layers which are groups of weights and parameters used by the neural network to generate its decisions.) obtaining sampling probability distribution and training iteration step arrangement, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution). Further Xiao, page 1229 -1230, section III – A, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement).) wherein the sampling probability distribution represents a probability that each of the M groups of parameters is sampled in each training iteration step, (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen) and the training iteration step arrangement comprises interval arrangement and periodic arrangement; (Xiao, page 1229 -1230, section III – A, and B, teaches the obtaining and use of the amount and structure of the epochs used in the training of the neural network models (i.e. training iteration step arrangement). This includes the determining of which layers should be frozen at each epoch.) freezing or stopping updating a sampled parameter group based on the sampling probability distribution and the training iteration step arrangement; (Xiao, page 1229 -1230, section III – B, C, and E, teach the calculating of a freezing rate that determines if the layers (i.e groups of parameters) gradients are likely to be canceled out and if they should be frozen at each epoch (i.e. sampling probability distribution) If the gradients are canceled out that means that the parameters are not being utilized in the epoch (i.e. training iteration step) and it is determined that the layer should be frozen and training the to-be-trained neural network based on the parameter group that is frozen or stopped updating. (Xiao, page 1229 -1230, section III – B, C, and E, teach the freezing of layers (i.e. parameter groups) in a neural network and then continuing training the neural network with those frozen layers) Regarding Claim 16 Xiao teaches The medium according to claim 15, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and freezing the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation and parameter update are not performed on the mth group of parameters to the first group of parameters. ; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen. When a layer is chosen to be frozen it is frozen for subsequent epochs and the gradient calculation and parameter updates are not performed on that layer.) Regarding Claim 17 Xiao teaches The medium according to claim 15, wherein the freezing or stopping updating the sampled parameter group based on the sampling probability distribution and the training iteration step arrangement comprises: determining a first iteration step based on the training iteration step arrangement, wherein the first iteration step is a to-be-sampled iteration step; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch) determining, based on the sampling probability distribution, an mth group of parameters sampled in the first iteration step, wherein m is a positive integer less than or equal to M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be frozen.) and stopping updating the mth group of parameters to a first group of parameters in the first iteration step, wherein the freezing the mth group of parameters to a first group of parameters in the first iteration step indicates that gradient calculation is performed and parameter update are not performed on the mth group of parameters to the first group of parameters. (Xiao, page 1229 -1230, section III – B, C, E, and D and Algorithm 1, Fig, $, teaches the analyzing of the layers and using a layer freezing rate for the current epoch (i.e. sample probability distribution) to determine if a layer (i.e. mth group of parameters) should be Stopped updating. In figure 4 it is shown that when a layer is frozen if the layer after it is unfrozen the gradient for the frozen layer will still be calculated but the parameters will not be updated (i.e. stopping updating).) Regarding Claim 18 Xiao teaches The medium according to claim 16, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 19 Xiao teaches The medium according to claim 17, wherein when in response to the training iteration step arrangement being the interval arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining a first interval; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. training iteration step arrangement) the first epoch is taken by the algorithm to determine what layers are to be frozen for the next epoch. The number of epochs before the next freezing occurs is determined and set as a hyperparameter in the freezing training algorithm and a first number of epochs is determined (i.e. first interval)) and determining one or more first iteration steps at every first interval in a plurality of training iteration steps. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of a number of epochs before each freeze (i.e. training interval) and an epoch that the freeze will be done on (i.e. iteration step), for every round of training that the model goes though (i.e. training iteration steps). Regarding Claim 20 Xiao teaches The medium according to claim 16, wherein in response to the training iteration step arrangement being the periodic arrangement, the determining the first iteration step based on the training iteration step arrangement comprises: determining that a quantity of first iteration steps is M-1; (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the use of epochs and a hyperparameter determine how many epochs will be done (i.e. iteration step arrangement))) and determining a first period based on the quantity of first iteration steps and a first proportion, wherein the first period comprises the first iteration step and an iteration step to be trained on the entire network, the first proportion is a proportion of the first iteration step in the first period, and the first iteration step is last (M-1) iteration steps in the first period. (Xiao, page 1229 -1230, section III – B, C, and E, and Algorithm 1, teaches the periodic freezing of layers based on how many epochs have been performed (i.e. iteration step) as the epochs are performed the freezing rate increases for each individual layer and when the set period of epochs is hit the layers indicated by the freezing rate will be frozen or stopped till the ending of the training epochs (i.e. last iteration step in the first period.) ) Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THOMAS B LANE whose telephone number is (571)272-1872. The examiner can normally be reached M-Th: 7:20am-5:20pm; F: Out of Office. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MARIELA REYES can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /THOMAS BERNARD LANE/Examiner, Art Unit 2142 /HAIMEI JIANG/Primary Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

May 23, 2023
Application Filed
Sep 29, 2023
Response after Non-Final Action
Mar 09, 2026
Non-Final Rejection mailed — §102
Jun 09, 2026
Response Filed
Sep 21, 2026
Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718117
DISTRIBUTED COMPUTING FOR DYNAMIC GENERATION OF OPTIMAL AND INTERPRETABLE PRESCRIPTIVE POLICIES WITH INTERDEPENDENT CONSTRAINTS
4y 10m to grant Granted Aug 25, 2026
Patent 12712765
METHOD AND DEVICE FOR ESTIMATING A CHANNEL, AND ASSOCIATED COMPUTER PROGRAM
3y 9m to grant Granted Aug 18, 2026
Patent 12705481
SEARCH METHOD, ELECTRONIC DEVICE AND STORAGE MEDIUM BASED ON NEURAL NETWORK MODEL
3y 11m to grant Granted Aug 11, 2026
Patent 12699880
DECENTRALIZED FEDERATED MACHINE-LEARNING BY SELECTING PARTICIPATING WORKER NODES
3y 7m to grant Granted Aug 04, 2026
Patent 12670386
MODEL TRAINING APPARATUS, MODEL TRAINING METHOD, AND COMPUTER-READABLE MEDIUM
4y 4m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
82%
With Interview (+5.0%)
3y 9m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 17 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month