Prosecution Insights
Last updated: October 02, 2026
Application No. 18/355,619

METHOD AND APPARATUS WITH MODEL TRAINING

Non-Final OA §101§102§103§112
Filed
Jul 20, 2023
Priority
Aug 15, 2022 — CN 202210975061.4 +1 more
Examiner
LAU, KAITLYN RENEE
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Non-Final)
60%
Grant Probability
Moderate
2-3
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 60% of resolved cases
60%
Career Allowance Rate
6 granted / 10 resolved
+5.0% vs TC avg
Strong +67% interview lift
Without
With
+66.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
27 currently pending
Career history
40
Total Applications
across all art units

Statute-Specific Performance

§101
28.8%
-11.2% vs TC avg
§103
34.3%
-5.7% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
21.9%
-18.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 10 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION This action is in response to the application filed 06/05/2026. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 Claims 4, 5, 14 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 4 recites the limitation "the second layers" in line 6. There is insufficient antecedent basis for this limitation in the claim. It is unclear as to whether these second layers are referring to a second layer of the plurality of layers or if these second layers are referring to something else. For purposes of examination, Examiner has determined the second layers to be the same as a second layer of the plurality of layers. Claim 5 recites the limitation “a maintenance probability value” in line 7. It is unclear as to whether this value is the same as the respective maintenance probability or if this value is a new value. For purposes of examination, Examiner has interpreted this maintenance probability value to be the same as the respective maintenance probability. Claim 14 recites the limitation "the second layers" in line 7. There is insufficient antecedent basis for this limitation in the claim. It is unclear as to whether these second layers are referring to a second layer of the plurality of layers or if these second layers are referring to something else. For purposes of examination, Examiner has determined the second layers to be the same as a second layer of the plurality of layers. Claim 15 recites the limitation “a maintenance probability value” in line 7. It is unclear as to whether this value is the same as the respective maintenance probability or if this value is a new value. For purposes of examination, Examiner has interpreted this maintenance probability value to be the same as the respective maintenance probability. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claims are directed towards an abstract idea without significantly more. Regarding Claim 1: Subject Matter Eligibility Analysis Step 1: Claim 1 recites a method and is thus a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 1 recites calculating a respective sensitivity of each layer of the plurality of layers (This limitation is a mental process as it encompasses a human mentally calculating a sensitivity and is thus an evaluation.) calculating a first maintenance probability for a t-th repeated training of the model; (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) calculating a respective maintenance probability of each of the plurality of layers based on the respective sensitivity of the plurality of layers and based on the first maintenance probability for the t-th repeated training of the model (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) Therefore, claim 1 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 1 further recites additional elements of A processor-implemented method (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) iteratively training a model through repeated training operations (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of iteratively training (see MPEP 2106.05(g)).) the model including a plurality of layers (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) performing the t-th repeated training of the model by training selected one or more maintenance layers selected from the plurality of layers, the one or more maintenance layers including respective maintenance probabilities satisfying a first predetermined maintenance condition. (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of iteratively training (see MPEP 2106.05(g)).) Therefore, claim 1 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 1 do not provide significantly more than the abstract idea itself, taken alone and in combination because A processor-implemented method uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). iteratively training a model through repeated training operations is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al., page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). the model including a plurality of layers uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). performing the t-th repeated training of the model by training selected one or more maintenance layers selected from the plurality of layers, the one or more maintenance layers including respective maintenance probabilities satisfying a first predetermined maintenance condition is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al., page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 1 is subject-matter ineligible. Regarding Claim 2: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 2 recites wherein, in the calculating of the respective sensitivity of each layer, a corresponding sensitivity of an 1-th layer is calculated based on a first accuracy of the model resulting from the plural layers being trained a predetermined number of times and a second accuracy of the model resulting from less than the plural layers, with training of the I-th layer being skipped, trained a corresponding predetermined number of times, and wherein "l" is a positive integer and has a value not greater than a number of layers of the model. (This limitation is a mental process as it encompasses a human mentally calculating a sensitivity and is thus an evaluation.) Therefore, claim 2 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 2 does not further recite any additional elements. Therefore, claim 2 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 2 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 2 is subject-matter ineligible. Regarding Claim 3: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 3 recites wherein the first maintenance probability of the t-th repeated training of the model is calculated based on a related parameter of the model, a training repetition ordinal number "t," and a predetermined maintenance probability, and wherein "t" is a positive integer. (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) Therefore, claim 3 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 3 does not further recite any additional elements. Therefore, claim 3 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 3 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 3 is subject-matter ineligible. Regarding Claim 4: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 4 recites determining first layers, of the plurality of layers, the first layers including a respective sensitivity satisfying a predetermined sensitivity condition as a maintained layer, of the one or more maintenance layers, that is to be maintained for each of plural repeated trainings; (This limitation is a mental process as it encompasses a human mentally determining layers that satisfy a condition and is thus an evaluation.) determining a second layer, of the plurality of layers, the second layers including a respective sensitivity satisfying a second predetermined sensitivity condition as a skipped layer for which training is to be skipped in each of the plural repeated trainings. (This limitation is a mental process as it encompasses a human mentally determining layers that satisfy a condition and is thus an evaluation.) Therefore, claim 4 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 4 does not further recite any additional elements. Therefore, claim 4 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 4 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 4 is subject-matter ineligible. Regarding Claim 5: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 5 recites calculating respective maintenance probabilities of each of one or more layers of the plurality of layers, other than the one or more maintenance layers and the skipped layer, for the t-th repeated training of the model; (This limitation is a mental process as it encompasses a human mentally calculating probabilities and is thus an evaluation.) setting the respective maintenance probability of each of the one or more maintenance layers with a maintenance probability value satisfying the first predetermined maintenance condition. (This limitation is a mental process as it encompasses a human mentally setting probabilities and is thus an judgment.) Therefore, claim 5 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 5 does not further recite any additional elements. Therefore, claim 5 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 5 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 5 is subject-matter ineligible. Regarding Claim 6: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 6 recites calculating a calibration factor of the t-th repeated training of the model, based on a current throughput of the model and the first maintenance probability of the t-th repeated training of the model; (This limitation is a mental process as it encompasses a human mentally calculating a calibration factor and is thus an evaluation.) calculating respective maintenance probabilities of each layer of the plurality of layers of the model based on the respective sensitivity of each of the plurality of layers of the model, the first maintenance probability of the t-th repeated training of the model, and the calibration factor of the t-th repeated training of the model. (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) Therefore, claim 6 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 6 does not further recite any additional elements. Therefore, claim 6 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 6 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 6 is subject-matter ineligible. Regarding Claim 7: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 7 recites calculating the first maintenance probability of the t-th repeated training of the model in accordance with: θ t = 2 ( a + c ) Γ ( a + c ) b ( a + c ) ( t - ε ) ( a + c - 1 ) e ( - 2 * t - ε b ) η θ 2 + θ and wherein θ t is the first maintenance probability of the t-th repeated training of the model, a is a shape parameter of the model, b is a proportional parameter of the model, c is a binomial weight of the model, t is a training repetition ordinal number, ε is a threshold parameter of the model, η is an amplification factor of the model, θ is a predetermined maintenance probability and Γ is a gamma function (This limitation is a mathematical concept as it encompasses a mathematical formula.) Therefore, claim 7 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 7 does not further recite any additional elements. Therefore, claim 7 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 7 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 7 is subject-matter ineligible. Regarding Claim 8: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 8 recites wherein the calculating of the respective maintenance probability of each of the plurality of layers of the model, based on the respective sensitivity of each of the plurality of layers of the model, the maintenance probability of the t-th repeated training of the model, and the calibration factor for the t-th repeated training of the model includes calculating the respective maintenance probability of each of the plural layers of the model in accordance with: p t , l = c l a m p ( α t ( θ t + β S b a s e ( l ) ,   θ m i n ,   θ m a x ) wherein pt,l is the respective maintenance probability of an l-th layer for the t-th repeated training of the model, α t is the calibration factor for the t-th repeated training of the model, θ t is the first maintenance probability for the t-th repeated training of the model, β is a sensitivity factor, S b a s e ( l ) is sensitivity of the l-th layer of the model, θ m i n is a minimum value for the respective maintenance probability of the l-th layer of the t-th repeated training of the model, and θ m a x is a maximum value of the respective maintenance probability of the l-th layer for the t-th repeated training of the model (This limitation is a mathematical concept as it encompasses a mathematical formula.) Therefore, claim 8 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 8 does not further recite any additional elements. Therefore, claim 8 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 8 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 8 is subject-matter ineligible. Regarding Claim 9: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 9 recites wherein the calculating of the calibration factor for the t-th repeated training of the model, based on the current throughput of the model and the first maintenance probability for the t-th repeated training of the model includes calculating the calibration factor for the t-th repeated training of the model in accordance with: α t = 2   -   ( T P c u r r θ t + x - θ t * x ) , and wherein α t is the calibration factor of the t=th repeated training of the model, T P c u r r is the current throughput of the model, θ t is the first maintenance probability for the t-th repeated training of the model, and x is a predetermined throughput improvement goal. (This limitation is a mathematical concept as it encompasses a mathematical formula.) Therefore, claim 9 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 9 does not further recite any additional elements. Therefore, claim 9 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 9 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 9 is subject-matter ineligible. Regarding Claim 10: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 10 recites determining whether an experiment result of a Bernoulli distribution including a respective third maintenance probability of each layer as a parameter is "l"; (This limitation is a mental process as it encompasses a human mentally determining whether a result is l and is thus an evaluation.) determining one or more layers having a Bernoulli distribution value corresponding to "I" as a maintenance layer of the one or more maintenance layers. (This limitation is a mental process as it encompasses a human mentally determining layers as maintenance layers and is thus an evaluation.) Therefore, claim 10 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 10 does not further recite any additional elements. Therefore, claim 10 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 10 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 10 is subject-matter ineligible. Regarding Claim 11: Subject Matter Eligibility Analysis Step 1: Claim 11 recites an apparatus and is thus a machine, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 11 recites calculate a respective sensitivity of layers included in a model, (This limitation is a mental process as it encompasses a human mentally calculating a sensitivity and is thus an evaluation.) calculate a first maintenance probability for a t-th repeated training of the model; (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) calculate a respective maintenance probability of layers of the model based on the respective sensitivity of the layers based on the first maintenance probability for the t-th repeated training of the model (This limitation is a mental process as it encompasses a human mentally calculating a probability and is thus an evaluation.) Therefore, claim 11 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 11 further recites additional elements of An electronic apparatus, the apparatus comprising: a processor configured to: (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) perform the t-th repeated training of the model by training selected one or more maintenance layers of the layers, the one or more maintenance layers selected from respective maintenance probabilities satisfying a first predetermined maintenance condition. (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of iteratively training (see MPEP 2106.05(g)).) Therefore, claim 11 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 11 do not provide significantly more than the abstract idea itself, taken alone and in combination because An electronic apparatus, the apparatus comprising: a processor configured to: uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). perform the t-th repeated training of the model including training selected one or more maintenance layers, of the plural layers of the model, whose respective maintenance probabilities satisfy a first predetermined maintenance condition is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al., page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 11 is subject-matter ineligible. Regarding claim 12, claim 12 recites substantially similar limitations to claim 2, and is therefore rejected under the same analysis. Regarding claim 13, claim 13 recites substantially similar limitations to claim 3, and is therefore rejected under the same analysis. Regarding claim 14, claim 14 recites substantially similar limitations to claim 4, and is therefore rejected under the same analysis. Regarding claim 15, claim 15 recites substantially similar limitations to claim 5, and is therefore rejected under the same analysis. Regarding claim 16, claim 16 recites substantially similar limitations to claim 6, and is therefore rejected under the same analysis. Regarding claim 17, claim 17 recites substantially similar limitations to claim 7, and is therefore rejected under the same analysis. Regarding claim 18, claim 18 recites substantially similar limitations to claim 8, and is therefore rejected under the same analysis. Regarding claim 19, claim 19 recites substantially similar limitations to claim 10, and is therefore rejected under the same analysis. Regarding Claim 20: Subject Matter Eligibility Analysis Step 1: Claim 20 recites a method and is thus a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 20 recites determining, from among a plurality of layers of a machine-learning model, one or more layers having a sensitivity below a predetermined threshold, wherein the sensitivity is evaluated based on a probability for a t-th repeated training of the machine-learning model; (This limitation is a mental process as it encompasses a human mentally determining a layer having a sensitivity below a threshold and is thus an evaluation.) Therefore, claim 20 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 20 further recites additional elements of iteratively training the machine-learning model as t-th repeated training, including skipping training of the one or more layers having the sensitivity below the predetermined threshold; (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) training the machine-learning model according to remaining layers, other than the one or more layers whose training is skipped in the t-th repeated training, having sensitivities above the predetermined threshold. (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 20 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 20 do not provide significantly more than the abstract idea itself, taken alone and in combination because iteratively training the machine-learning model as t-th repeated training, including skipping training of the one or more layers having the sensitivity below the predetermined threshold uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). training the machine-learning model according to remaining layers, other than the one or more layers whose training is skipped in the t-th repeated training, having sensitivities above the predetermined threshold uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 20 is subject-matter ineligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-5, 11-15, and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Bijalwan et al. (US 2023/0281423 A1) (hereafter referred to as Bijalwan). Regarding claim 1, Bijalwan teaches A processor-implemented method, the method comprising: iteratively training a model through repeated training operations, the model including a plurality of layers, the iterative training comprising (Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the repeated training steps are the repeated steps from 314 back to 308 and 322 back to 308 of Figure 3.): calculating a respective sensitivity of each layer of the plurality of layers (Bijalwan, page 13, paragraph 0036, “In one embodiment, the sensitivity evaluation module 114 generates a base model from the input neural network model by representing the parameters of the input neural network model in high precision format and stores as base model 210. The sensitivity evaluation module 114 also generates a plurality of weight evaluation models 212 for each layer of the input neural network model to evaluate a weight sensitivity value. The sensitivity evaluation module 114 generates a plurality of feature evaluation models 216 for each layer of the input neural network to evaluate a feature sensitivity value. The sensitivity evaluation module generates a union sensitivity list 220 based on the weight sensitivity values and feature sensitivity values evaluated for each layer of the input neural network model.” Examiner notes that the neural network is a machine learning model. Examiner further notes that a sensitivity value is evaluated or calculated for each layer in the model.); calculating a first maintenance probability for a t-th repeated training of the model (Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first maintenance probability is the second accuracy. Examiner further notes that t-th repeated training is the repeated steps from 314 back to 308 of Figure 3.); calculating a respective maintenance probability of each of the plurality of layers based on the respective sensitivity of the plurality of layers and based on the first maintenance probability for the t-th repeated training of the model (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first accuracy is the respective maintenance probability of each of the plurality layers. Examiner further notes that the respective maintenance probability is based on the first maintenance probability through the loop of step 308 through step 326 of Figure 3.); and performing the t-th repeated training of the model by training one or more maintenance layers selected from the plurality of layers, the one or more maintenance layers including respective maintenance probabilities satisfying a first predetermined maintenance condition (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the selected one or more maintenance layers is the group of layers. Examiner further notes that the first predetermined maintenance condition is the threshold as shown in step 314.). Regarding claim 2, Bijalwan teaches The method of claim 1, wherein, in the calculating of the respective sensitivity of each layer included in the model, a corresponding sensitivity of an l-th layer is calculated based on a first accuracy of the model resulting from the plural layers being trained a predetermined number of times and a second accuracy of the model resulting from less than the plural layers, (Bijalwan, page 15, paragraph 0052, “The sensitivity evaluation module 114 calculates a base model weight output of a final layer of the base model. The sensitivity evaluation module 114 calculates a first difference metric between the first weight output and the base model weight output and a second difference metric between the second weight output and the base model weight output. In one embodiment, the difference metric may be calculated using a Jacobian determinant standard deviation method. The sensitivity evaluation module 114 evaluates a mean of the first difference metric and the second difference metric and stores the value as the weight sensitivity value of the layer” where “the present disclosure provides an efficient method quantization aware training of a neural network model by quantizing a group of layers in a single iteration and validating the model” (Bijalwan, page 16, paragraph 0072). Examiner notes that the base model weight output is the first and second accuracy. Examiner further notes that the predetermined number of times the layers are trained is once. Additionally, Examiner notes that the accuracy of the base model is a result of the final layer which is less than the total amount of layers.) with training of the l-th layer being skipped, trained a corresponding predetermined number of times and wherein "l" is a positive integer and has a value not greater than a number of layers of the model (Bijalwan, page 12, paragraph 0025, “Embodiments of the present disclosure relates to an efficient method of quantization aware training of a neural network model by quantizing a group of layers in a single iteration and validating the model. Further, the present disclosure also provides an efficient method of grouping the layers based on their sensitivity values and quantizes the group of layers that correspond to highest sensitivity first to achieve the target accuracy. The present disclosure also achieves the target accuracy in less time by quantizing the group of layers that corresponding to high sensitivity first and less sensitivity later. Thus, the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy.” Examiner notes that ignoring quantization of layers is skipping training of the l-th layer. Examiner further notes that the single iteration is the predetermined number of times the layer is trained.). Regarding claim 3, Bijalwan teaches The method of claim 1, wherein the first maintenance probability of the t-th repeated training of the model is calculated based on a related parameter of the model, a training repetition ordinal number "t," and a predetermined maintenance probability, and wherein "t" is a positive integer (Bijalwan, page 14, paragraph 0039, “The quantization module 118 may make a forward pass of the temporary model and may calculate an accuracy of the temporary model. The quantization module 118 may compare the calculated accuracy of the temporary model with a known target accuracy. The quantization module 118 may retain the quantization if the quantization results in significant improvement to achieve target accuracy. The quantization module 118 may reject the quantization if the quantization does not result in significant accuracy improvement” where “The data acquisition module 230 may receive information about the neural network model such as number of layers in the neural network model, one or more parameters such as, but not limited to, inputs, weights, biases, activation functions, outputs associated with each layer” (Bijalwan, page 13, paragraph 0035) and where “The quantization module 118 selects the first group in a first iteration, the second group in a second iteration and the third group in a third iteration” (Bijalwan, page 16, paragraph 0063) and Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first maintenance probability is the second accuracy. Examiner further notes that t-th repeated training is the repeated steps from 314 back to 308 of Figure 3. Examiner additionally notes that the parameters that are received from the data acquisition module are the related parameters, and the three iterations is the ordinal number t. Examiner also notes that the predetermined maintenance probability is the target accuracy in step 322 of Figure 3.). Regarding claim 4, Bijalwan teaches The method of claim 1, further comprising: determining first layers, of the plurality of layers, the first layers including a respective sensitivity satisfying a predetermined sensitivity condition as a maintained layer, of the one or more maintenance layers, to be maintained for each of plural repeated trainings (Bijalwan, page 16, paragraph 0062, “At block 506, the grouping module 116 clusters the layers into groups based on the plurality of thresholds computed at block 504. The grouping module 116 clusters first set of sensitivity values associated with scores of greater than or equal to the first threshold into a first group. The grouping module 116 clusters a second set of sensitivity values associated with scores of greater than the second threshold and less than the first threshold into a second group. The grouping module 116 clusters a third set of sensitivity values associated with scores of greater than the third threshold and less than the second threshold into a third group. The grouping module 116 clusters the remaining set of sensitivity values into a fourth group” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025). Examiner notes that quantizing the layers that contribute to more loss or that are more sensitive is maintaining the layers that satisfy the predetermined sensitivity condition of the thresholds.); and determining a second layer, of the plurality of layers, the second layers including a respective sensitivity satisfying a second predetermined sensitivity condition as a skipped layer for which training is to be skipped in each of the plural repeated trainings(Bijalwan, page 16, paragraph 0062, “At block 506, the grouping module 116 clusters the layers into groups based on the plurality of thresholds computed at block 504. The grouping module 116 clusters first set of sensitivity values associated with scores of greater than or equal to the first threshold into a first group. The grouping module 116 clusters a second set of sensitivity values associated with scores of greater than the second threshold and less than the first threshold into a second group. The grouping module 116 clusters a third set of sensitivity values associated with scores of greater than the third threshold and less than the second threshold into a third group. The grouping module 116 clusters the remaining set of sensitivity values into a fourth group” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025). Examiner notes ignoring quantization of the layers that negligible improvement is skipping layers whose sensitivity satisfy the second predetermined condition of not providing enough improvement.). Regarding claim 5, Bijalwan teaches The method of claim 4, wherein the calculating of the respective maintenance probability of each of the plurality of layers comprises: calculating respective maintenance probabilities of each of one or more layers of the plurality layers, other than the one or more maintenance layers and the skipped layer, for the t-th repeated training of the model (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025) and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first accuracy is the respective maintenance probability. Examiner further notes that the first accuracy is only calculated for one group of layers at a time. The layers that are ignored or already selected as maintenance layers are not calculated.); and setting the respective maintenance probability of each of the one or more maintenance layers with a maintenance probability value that satisfies the first predetermined maintenance condition (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025) and Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first accuracy is the respective maintenance probability. Examiner further notes that the maintenance probability is also the first accuracy and the first predetermined maintenance condition is the threshold 314 of Figure 3.). Regarding claim 11, Bijalwan teaches An electronic apparatus, the apparatus comprising: a processor configured to: (Bijalwan, page 11, paragraph 0009, “The system comprises a memory and a processor that is coupled to the memory.”.): calculate a respective sensitivity of layers in a model (Bijalwan, page 13, paragraph 0036, “In one embodiment, the sensitivity evaluation module 114 generates a base model from the input neural network model by representing the parameters of the input neural network model in high precision format and stores as base model 210. The sensitivity evaluation module 114 also generates a plurality of weight evaluation models 212 for each layer of the input neural network model to evaluate a weight sensitivity value. The sensitivity evaluation module 114 generates a plurality of feature evaluation models 216 for each layer of the input neural network to evaluate a feature sensitivity value. The sensitivity evaluation module generates a union sensitivity list 220 based on the weight sensitivity values and feature sensitivity values evaluated for each layer of the input neural network model.” Examiner notes that the neural network is a machine learning model. Examiner further notes that a sensitivity value is evaluated or calculated for each layer in the model.); calculate a first maintenance probability for a t-th repeated training of the model (Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first maintenance probability is the second accuracy. Examiner further notes that t-th repeated training is the repeated steps from 314 back to 308 of Figure 3.); calculate a respective maintenance probability of the layers of the model based on the respective sensitivity of the layers and based on the first maintenance probability for the t-th repeated training of the model (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first accuracy is the respective maintenance probability of each of the plural layers. Examiner further notes that the respective maintenance probability is based on the first maintenance probability through the loop of step 308 through step 326 of Figure 3.); and perform the t-th repeated training of the model by training one or more maintenance layers of the layers, the one or more maintenance layers selected from respective maintenance probabilities satisfying a first predetermined maintenance condition (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the selected one or more maintenance layers is the group of layers. Examiner further notes that the first predetermined maintenance condition is the threshold as shown in step 314.). Regarding claim 12, claim 12 recites substantially similar limitations to claim 2, and is therefore rejected under the same analysis. Regarding claim 13, claim 13 recites substantially similar limitations to claim 3, and is therefore rejected under the same analysis. Regarding claim 14, claim 14 recites substantially similar limitations to claim 4, and is therefore rejected under the same analysis. Regarding claim 15, claim 15 recites substantially similar limitations to claim 5, and is therefore rejected under the same analysis. Regarding claim 20, Bijalwan teaches A processor-implemented method, the method comprising: determining from among a plurality of layers of a machine-learning model, one or more layers having sensitivity below a predetermined threshold, wherein the sensitivity is evaluated based on a probability for a t-th repeated training of the machine learning model (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 50% and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025) and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the predetermined threshold is the threshold as shown in step 314.); Iteratively training the machine-learning model as t-th repeated training, including skipping training of the one or more layers having the sensitivity below the predetermined threshold (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 50% and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025) and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes ignoring quantization of layers that do not provide significant improvement, is skipping training of the layers having the sensitivity below the predetermined threshold where the threshold is the threshold in step 314 of Figure 3.); And training the machine-learning model according to remaining layers, other than the one or more layers whose training is skipped in the t-th repeated training, having sensitivities above the predetermined threshold (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 50% and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” where “the present disclosure reduces or limits the training time of the neural network model by quantizing only those layers that contribute to more loss or that are more sensitive and that contribute significant improvement in accuracy and ignoring quantization of those layers that provide negligible improvement in accuracy” (Bijalwan, page 12, paragraph 0025) and Bijalwan, page 4, Figure 3 PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes ignoring quantization of layers that do not provide significant improvement, is skipping training of the layers having the sensitivity below the predetermined threshold where the threshold is the threshold in step 314 of Figure 3. Examiner further notes that layers having high sensitivities meet the threshold in step 314 and are processed more.). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 6, 10, 16, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan in view of Zhang et al. (“Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping”). Regarding claim 6, Bijalwan teaches the method of claim 1. Bijalwan further teaches calculating respective maintenance probabilities of each layer of the plurality of layers based on the respective sensitivity of each of the plurality of layers of the model, the first maintenance probability of the t-th repeated training of the model, …of the t-th repeated training of the model (Bijalwan, page 17, paragraph 0082-0083, The MPQS 102 clusters layers into a number of groups using the union sensitivity list. The MPQS 102 selects a first group of layers with high sensitivity and quantizes the group into a high precision format as described above to generate a temporary model. For example, the temporary model comprises quantization of weights of layer 1 and features of layer 5 into high precision format. [0083] The MPQS 102 makes a forward pass including the quantization within the temporary model and computes a first accuracy of the temporary model. The MPQS 102 compares the first accuracy, for e.g., 60% with the threshold accuracy i.e., 500/o and retains the quantization since the first accuracy is greater than the threshold accuracy. The MPQS 102 retains the quantization of the temporary model by including the quantization of the temporary model to the input neural network model and storing it as a mixed precision model. The MPQS 102 computes a second accuracy, also called "accuracy", of the mixed precision model with the target accuracy. If the accuracy is 65%<target accuracy 90%, the MPQS 102 proceeds to quantize a next group of layers. The MPQS generates another temporary model and computes a first accuracy of another temporary model. For e.g., another temporary model comprises quantizing weights of 6th and 7th layers and features of 11 th and 15th layer into high precision format. If the first accuracy is 70%, the MPQS 102 updates the mixed precision model previously stored with the quantization of another temporary model. The MPQS 102 then computes a second accuracy for e.g., 95% and compares with target accuracy 90%. The MPQS 102 stores the updated mixed precision model as the final mixed precision model of the input neural network model” and Bijalwan, page 4, Figure 3 [see below] PNG media_image1.png 991 978 media_image1.png Greyscale Examiner notes that the first accuracy is the respective maintenance probability of each of the plurality layers. Examiner further notes that the respective maintenance probability is based on the first maintenance probability through the loop of step 308 through step 326 of Figure 3.) Bijalwan does not explicitly teach a calibration factor. However, Zhang teaches wherein the calculating of the respective maintenance probability of each of the plurality of layers comprises: calculating a calibration factor of the t-th repeated training of the model, based on a current throughput of the model and the first maintenance probability of the t-th repeated training of the model (Zhang, page 5, 4th paragraph - 5th paragraph, “Next, we extend the architecture to include a gate for each sub-layer (Fig. 5c), which controls whether a sub-layer is disabled or not during training. In particular, for each mini-batch, the two gates for the two sublayers decide whether to remove their corresponding transformation functions and only keep the identify mapping connection, which is equivalent to applying a conditional gate function G to each sub-layer as follows…In our design, the function Gi only takes 0 or 1 as values, which is chosen randomly from a Bernoulli distribution (with two possible outcomes), Gi ~ B(1, pi), where pi is the probability of choosing 1. Because the blocks are selected with probability pi during training and are always presented during inference, we re-calibrate the layers’ output by a scaling factor of 1/pi whenever they are selected” and Zhang, page 6, Algorithm 1 PNG media_image2.png 376 312 media_image2.png Greyscale Examiner notes that the scaling factor is the calibration factor, θ - is the first maintenance probability and θ t is the respective maintenance probability.); and calculating respective maintenance probabilities of each layer of the plurality of layers based on … the first maintenance probability of the t-th repeated training of the model, and the calibration factor of the t-th repeated training of the model (Zhang, page 5, 4th paragraph - 5th paragraph, “Next, we extend the architecture to include a gate for each sub-layer (Fig. 5c), which controls whether a sub-layer is disabled or not during training. In particular, for each mini-batch, the two gates for the two sublayers decide whether to remove their corresponding transformation functions and only keep the identify mapping connection, which is equivalent to applying a conditional gate function G to each sub-layer as follows…In our design, the function Gi only takes 0 or 1 as values, which is chosen randomly from a Bernoulli distribution (with two possible outcomes), Gi ~ B(1, pi), where pi is the probability of choosing 1. Because the blocks are selected with probability pi during training and are always presented during inference, we re-calibrate the layers’ output by a scaling factor of 1/pi whenever they are selected” and Zhang, page 6, Algorithm 1 PNG media_image2.png 376 312 media_image2.png Greyscale Examiner notes that the scaling factor is the calibration factor, θ - is the first maintenance probability and θ t is the respective maintenance probability. ) Bijalwan and Zhang are considered analogous to the claimed invention because they skip training on layers of machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Bijalwan to include a calibration factor like Zhang. Doing so is advantageous because it “achieves competitive performance to training a deep model from scratch at a faster rate” (Zhang, page 9, Conclusion). Regarding claim 10, Bijalwan teaches the method of claim 1. Bijalwan does not teach, but Zhang does teach further comprising selecting , comprising: determining whether an experiment result of a Bernoulli distribution including a respective third maintenance probability of each layer as a parameter is “1”; and determining one or more layers having a Bernoulli distribution value corresponding to "1" as a maintenance layer of the one or more maintenance layers (Zhang, page 5, 4th paragraph - 5th paragraph, “Next, we extend the architecture to include a gate for each sub-layer (Fig. 5c), which controls whether a sub-layer is disabled or not during training. In particular, for each mini-batch, the two gates for the two sublayers decide whether to remove their corresponding transformation functions and only keep the identify mapping connection, which is equivalent to applying a conditional gate function G to each sub-layer as follows…In our design, the function Gi only takes 0 or 1 as values, which is chosen randomly from a Bernoulli distribution (with two possible outcomes), Gi ~ B(1, pi), where pi is the probability of choosing 1.”). Bijalwan and Zhang are considered analogous to the claimed invention because they skip training on layers of machine learning models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Bijalwan to use a Bernoulli distribution like in Zhang. Doing so is advantageous because it “achieves competitive performance to training a deep model from scratch at a faster rate” (Zhang, page 9, Conclusion). Regarding claim 16, claim 16 recites substantially similar limitations to claim 6, and is therefore rejected under the same analysis. Regarding claim 19, claim 19 recites substantially similar limitations to claim 10, and is therefore rejected under the same analysis. Allowable Subject Matter Claims 7-9 and 17-18 would be allowable over the prior art of record if the 101 and 112(b) rejections are overcome and if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Specifically, regarding claim 7, “calculating the first maintenance probability of the t-th repeated training of the model in accordance with: θ t = 2 ( a + c ) Γ ( a + c ) b ( a + c ) ( t - ε ) ( a + c - 1 ) e ( - 2 * t - ε b ) η θ 2 + θ and wherein θ t is the first maintenance probability of the t-th repeated training of the model, a is a shape parameter of the model, b is a proportional parameter of the model, c is a binomial weight of the model, t is a training repetition ordinal number, ε is a threshold parameter of the model, η is an amplification factor of the model, θ is a predetermined maintenance probability and Γ is a gamma function” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Bijalwan and Nkemnole et al. (“Poly-Weighted Exponentiated Gamma Distribution with Application”) (hereafter referred to as Nkemnole). Bijalwan discloses the first maintenance probability (Bijalwan, page 4, Figure 3), a training repetition ordinal number (Bijalwan, page 14, paragraph 0039; page 13, paragraph 0035; page 16, paragraph 0063; page 4, Figure 3), a threshold (Bijalwan, page 17, paragraph 0082-0083; Bijalwan, page 4, Figure 3), a predetermined maintenance probability (Bijalwan, page 14, paragraph 0039; Bijalwan, page 13, paragraph 0035, page 16, paragraph 0063; page 4, Figure 3). Bijalwan fails to disclose a shape parameter, a proportional parameter, a binomial weight, an amplification factor, a gamma function, and the overall formula. Nkemnole discloses the gamma function formula (Nkemnole, page 3, Equation 12), but does not disclose an amplification factor nor a predetermined maintenance probability. Additionally, Nkemnole is not in the realm of machine learning or artificial intelligence and thus cannot be reasonably combined with other prior art to teach claim 7. Therefore, the prior art of record, individually, or in combination, does not disclose claim 7 as a whole. Specifically, regarding claim 8, “calculating the respective maintenance probability of each of the plural layers of the model in accordance with: p t , l = c l a m p ( α t ( θ t + β S b a s e ( l ) ,   θ m i n ,   θ m a x ) wherein pt,l is the respective maintenance probability of an l-th layer for the t-th repeated training of the model, α t is the calibration factor for the t-th repeated training of the model, θ t is the first maintenance probability for the t-th repeated training of the model, β is a sensitivity factor, S b a s e ( l ) is sensitivity of the l-th layer of the model, θ m i n is a minimum value for the respective maintenance probability of the l-th layer of the t-th repeated training of the model, and θ m a x is a maximum value of the respective maintenance probability of the l-th layer for the t-th repeated training of the model” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Bijalwan and Zhang. Bijalwan discloses the respective maintenance probability (Bijalwan, page 17, paragraph 0082-0083; Bijalwan, page 4, Figure 3), the first maintenance probability (Bijalwan, page 4, Figure 3), and the sensitivity of the l-th layer (Bijalwan, page 13, paragraph 0036). Bijalwan fails to disclose, the calibration factor, a sensitivity factor, a minimum value, a maximum value, and the overall formula. Zhang discloses the calibration factor (Zhang, page 5, 4th paragraph - 5th paragraph), but fails to disclose a sensitivity factor, a minimum value, a maximum value, and the overall formula. Therefore, the prior art of record, individually, or in combination, does not disclose claim 8 as a whole. Specifically, regarding claim 9, “calculating the calibration factor for the t-th repeated training of the model in accordance with: α t = 2   -   ( T P c u r r θ t + x - θ t * x ) , and wherein α t is the calibration factor of the t=th repeated training of the model, T P c u r r is the current throughput of the model, θ t is the first maintenance probability for the t-th repeated training of the model, and x is a predetermined throughput improvement goal” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Zhang. Zhang discloses the calibration factor, the throughput, and the first maintenance probability (Zhang, page 5, 4th paragraph - 5th paragraph; Zhang, page 6, Algorithm 1). Zhang fails to disclose a predetermined throughput improvement goal and the overall formula. Therefore, the prior art of record, individually, or in combination, does not disclose claim 9 as a whole. Claim 17 recites substantially similar limitations as claim 7 and is therefore allowable under the same rationale if the 101 rejections are overcome and if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 18 recites substantially similar limitations as claim 8 and is therefore allowable under the same rationale if the 101 and 112(b) rejections are overcome and if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Response to Arguments Examiner notes that the objections to the specification have been overcome in light of the instant amendments. Examiner notes that the objections to the claims have been overcome in light of the instant amendments. Examiner notes that most of the 112(b) rejections have been overcome in light of the instant amendments and arguments. Claims 5 and 15 have maintained the 112(b) rejection. On page 7, Applicant argues: Thus, example 39 is relevant to the present claims as an example of claimed training that is not abstract. For example 39, the Office states that the "limitation does not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols." Similarly, in claim 1, the following is recited: calculating a respective sensitivity of each layer, calculating a first maintenance probability for at-th repeated training, calculating a respective maintenance probability of each of the plurality of layers, and performing the t-th repeated training by training selected maintenance layers. The Applicant submits that this limitation is likewise non-abstract and "does not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols." The Applicant submits that the present claims are similar to those in Example 39. Thus, in the present claims, as in Example 39, these steps cannot be performed in the human mind. Furthermore, the claims do not recite a mathematical concept. As explained in MPEP 2106.04(a)(2), a claim does not recite a mathematical concept if it is only based on or involves a mathematical concept. The claims here do not recite any mathematical formula, equation, or calculation as such. Thus, unlike applying a mathematical formula, the Office has, at best, identified claims that "involve a mathematical concept." Accordingly, in the present claims, Step 2A is satisfied and claim 1, for example, recites statutory subject matter and thus, the dependent claims depending therefrom are likewise not abstract. Regarding the Applicant’s argument that claim 1 does not recite an abstract idea, Examiner respectfully disagrees. Specifically, calculating a respective sensitivity of each layer, calculating a first maintenance probability for a t-th repeated training and calculating a respective maintenance probability of each of the plurality of layers are mental processes in which a human can mentally compute or calculate a sensitivity, a first maintenance probability, and respective maintenance probability. On page 8, Applicant argues: That is, claim 1, as an example, provides an improvement by performing iterative training of a neural network through a specific process of calculating per-layer sensitivity, determining maintenance probabilities for each t-th repeated training iteration, selecting maintenance layers based on those probabilities, and training a curated set of layers in each iteration while skipping others. Furthermore, as a result of this novel training method, example neural networks may be improved. Thus, when viewed as a whole, the claim integrates its operations into a practical application. This selective training approach significantly reduces the computational cost and training time compared to conventional methods that train all layers in every iteration, while still preserving model accuracy. For example, paragraph [0048] provides: In an example, a model training method may reduce the negative effect of skip layer calculation on the overall training accuracy by fitting the effect of a skipped training layer on a model convergence in a training time dimension and determining a training repetition maintenance probability using a parameter related to a layer of the skipped training of the model. That is, examples of claim 1 provide an improvement in the functioning of a neural network and is thus an improvement in a technological field. Regarding the Applicant’s argument that these elements provide an improvement, Examiner respectfully disagrees. Specifically, Examiner notes that the stated skip layer calculation, model convergence, and time dimension from paragraph 0048 of the specification are not reflected in claim 1 and thus does not reflect the improvement stated in paragraph 0048 (MPEP 2106.04(d)(1)). On pages 17-19, Applicant argues: Second, the Office has incorrectly identified a technology that is considered "routine" or "well understood." That is, the present claims are clearly directed to technological improvements the Office has recognized as important as discussed above in The Reminder Memo. … Here, the claims provide an improvement in a technology - the operation of neural networks during iterative training and are thus operations that "are able in combination to perform functions that are not merely generic." That is, the present claims are not merely claiming the performance of well-known actions, such as performing routine calculations or storing data. Finally, the Office cites to MPEP 2106.05(f) to support its assertions. However, the Office's assertions are incorrect. For example, on page 10 of the Office Action, the Office asserts: • performing the t-th repeated training of the model including training selected one or more maintenance layers, of the plural layers of the model, whose respective maintenance probabilities satisfy a first predetermined maintenance condition is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A 1, Metzler et al., page 11, paragraph 0052, "a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position"). However, this assertion is unclear and ultimately incorrect. As discussed above, the MPEP provides examples of what is "well understood, routine, and conventional activity" and training of a model is not included. In addition, it is unclear what "Metzler" refers to in the context of Section 101 Prong 2 and Step 2B interpretation. "Metzler" does not appear in MPEP 2106. Nor do the Examiner or the Reminder Memo refer to "Metzler". Nonetheless, it appears that the Office is citing another patent application for the incorrect assertion that training is conventional activity. However, as discussed above and below, this is clearly incorrect. Indeed, simply because one application describes a conventional process does not result in all similar process types being anticipated or obvious, let alone being considered "well understood" as defined within MPEP 2106. Regarding the Applicant’s argument that iteratively training a model is not well-understood, routine, and conventional, Examiner respectfully disagrees. Specifically, Examiner notes that MPEP section 2106.05(d)(I)(2) states “a factual determination is required to support a conclusion that an additional element (or combination of elements) is well-understood, routine, conventional activity” where “the required factual determination must be expressly supported in writing as discussed in MPEP §2106.07(a). Appropriate forms of support include one or more of the following: … (c) A citation to a publication that demonstrates the well-understood, routine, conventional nature of the additional elements(s).” As such, Metzler is referring to a publication which states that a iterative machine learning training process is conventional, and thus, performing t-th repeated training is well-known, routine, conventional as well. On page 20, Applicant argues: Finally, the Office cites to MPEP 2106.05(f) in general without specifically citing to any particular guidance from this section to support the assertion that "uses a computer as a tool to perform the abstract idea and cannot provide significantly more." However, the Applicant submits that MPEP 2106.05(f) is being improperly applied to the current claims. For example, MPEP 2106.05(f) provides examples for whether the claims are "an apply it" type of claim: Other examples where the courts have found the additional elements to be mere instructions to apply an exception, because they recite no more than an idea of a solution or outcome include: i. Remotely accessing user-specific information through a mobile interface and pointers to retrieve the information without any description of how the mobile interface and pointers accomplish the result of retrieving previously inaccessible information, Intellectual Ventures v. Erie lndem. Co., 850 F.3d 1315, 1331, 121 USPQ2d 1928, 1939 (Fed. Cir. 2017); ii. A general method of screening emails on a generic computer without any limitations that addressed the issues of shrinking the protection gap and mooting the volume problem, Intellectual Ventures Iv. Symantec Corp., 838 F.3d 1307, 1319, 120 USPQ2d 1353, 1361 (Fed. Cir. 2016); and iii. Wireless delivery of out-of-region broadcasting content to a cellular telephone via a network without any details of how the delivery is accomplished, Affinity Labs of Texas v. DirecTV, LLC, 838 F.3d 1253, 1262-63, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016). That is, these examples illustrate what 2106.05(f) is directed to, general ideas of performing an exception without specific guidance outside of an idea of a solution. On the other hand, the present claims specifically provide instructions for performing the claimed training. Therefore, the Office has incorrectly applied an "apply-it" style rejection. Regarding the Applicant’s argument that 2106.05(f) is being improperly applied, Examiner respectfully disagrees. Specifically, Examiner notes that the limitations in claim 1 which fall under 2106.05(f) are “A processor-implemented method” and “the model including a plurality of layers.” Applicant is only citing one of the three reasons an examiner would consider a limitation as mere instructions to implement an abstract idea on an computer. MPEP 2106.05(f) states: “When determining whether a claim simply recites a judicial exception with the words ‘apply it’ (or an equivalent), such as mere instructions to implement an abstract idea on a computer, examiners may consider the following: (1) Whether the claim recites only the idea of a solution outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished…. (2) Whether the claim invokes computers or other machinery merely as a tool to perform an existing process….(3) The particularity or generality of the application of the judicial exception.” In the case of claim 1, the limitations fall under (2) Whether the claim invokes computers or other machinery merely as a tool to perform an existing process, not (1) Whether the claim recites only the idea of a solution or outcome like the Applicant is claiming. On pages 22-24, Applicant argues: However, the cited portions of Bijalwan describe a quantization process in what Bijalwan describes as quantization-aware training. Thus, in Bijalwan, a group of layers is selected and quantized into a high precision format. A temporary model is created for the quantized high precision values and then an accuracy of that temporary model is measured. If its accuracy meets a certain level, it's stored as a mixed-precision model. Essentially, in Bijalwan, weights for different layers are quantized and their performances are assessed. That is, Bijalwan is silent with respect to training the model as recited, for example, in claim 1. Bijalwan uses sensitivity to decide how precisely to represent already trained layers. Examples of claim 1 use sensitivity to decide whether to train certain layers in the first place. Thus, in FIG. 3 of Bijalwan, the relevant steps towards Bijalwan's quantization-aware training is receiving a model and generating the union sensitivity list in step 304. A temporary model is created based on this list in Step 310 "by quantizing all the layers corresponding to the selected group into high precision based on the type of sensitivity value". Bijalwan ultimately provides a mixed-sensitivity model by relying on quantizing highest sensitivity layers first. Bijalwan's use of the term "Quantization-Aware Training" (QAT) is its own internal terminology for optimizing and "training" the quantization levels and precision assignments for a received, trained model. It is not actual iterative training of the neural network weights. Bijalwan's process is a quantization configuration search separate from a training of the obtained model. Thus, Bijalwan is so far conceptually removed from the present claims that there is no manner in which one of ordinary skill in the art would read Bijalwan and arrive at the present claims. On the other hand, in the present claims, the sensitivity is based on an impact of the layer when not trained in a current training iteration. Bijalwan uses its concept of precision to assess how sensitive a layer is to quantization. That is, how a layer reacts to a change in precision for that weight. Thus, in Bijalwan, the layers may be grouped together for different quantization inputs of varying precision. However, Bijalwan fails to describe or suggest any element of claim 1 as they are conceptually unrelated. Finally, while Bijalwan occasionally uses the word "training" in its disclosure, this refers only to forward passes performed on temporary quantized models post-training in a quantization optimization process. That is, where Bijalwan uses the term "training," Bijalwan is not referring to the claimed iterative training of the neural network weights. Instead, Bijalwan describes a "quantization aware training of a neural network for image compression" Thus, in paragraph [0030] of Bijalwan, a training database of images is referred to and, more specifically, paragraph [0038] of Bijalwan further explains (with emphasis added): ... Quantization of the neural network model may be performed using two techniques-Post-training quantization and Quantization-aware training. Post training quantization is a technique in which the neural network is trained using floating-point computation and then quantized after the training. Quantization aware training generates a quantized version of the neural network in a forward pass and parallelly trains the neural network using the quantized version. Methods in the present disclosure preferably employ Quantization-aware training technique. Bijalwan's quantization-aware training thus refers to quantizing the neural network to customize its quantization levels. Bijalwan generates a quantized version of a temporary model in a forward pass and then evaluates that quantized version to determine optimal quantization parameters. Thus, while paragraphs [0072 and 0082] of Bijalwan describe a "method quantization aware training of a neural network model by quantizing a group of layers in a single iteration and validating the model," Bijalwan explains, at paragraph [0051 ], this as receiving a neural network model and generating a base model from that model by "quantizing all the parameters of the neural network model into high precision format" and then: At block 404, the sensitivity evaluation module 114 evaluates a weight sensitivity value for each parametric layer. The sensitivity evaluation module 114 generates the first weight evaluation model and the second evaluation model and stores them as the weight evaluation models 212 for the layer. The sensitivity evaluation module 114 compares outputs of the first weight evaluation model and the second weight evaluation model with the base model to compute a first and a second weight sensitivity values respectively. The sensitivity evaluation module 114 determines a mean of the first and second weight sensitivity values as the weight sensitivity value of the parametric layer and stores the weight sensitivity value in the weight sensitivity data 214. Thus, while Bijalwan refers to its process as "quantization-aware training," this terminology is misleading. Bijalwan obtains a pre-trained neural network, generates a union sensitivity list, creates temporary quantized models, and evaluates those temporary models to produce a final mixed-precision model. In reality, Bijalwan's so-called "training" is actually a quantization optimization performed on an already-trained model - not actual iterative training of the neural network weights. Regarding the Applicant’s argument that Bijalwan does not teach training the model, Examiner respectfully disagrees. Specifically, Examiner notes that claim 1 recites “iteratively training a model through repeated training operations” in which these training operations include calculating sensitivities and probabilities. The only limitation that further describes this iterative training is “performing the t-th repeated training of the model by training one or more maintenance layers selected from the plurality of layers, the one or more maintenance layers including respective maintenance probabilities satisfying a first predetermined maintenance condition.” Under broadest reasonable interpretation, performing the t-th repeated training of the model by training one or more maintenance layers encompasses repeating the steps of calculated sensitivities and probabilities on selected layers. Bijalwan demonstrates this in Figure 3 and paragraphs 0082-0083. Examiner respectfully directs the Applicant to the above prior art rejections. Thus, Bijalwan’s Quantization-Aware Training maps onto claim 1. On page 16, Applicant argues: The rejected dependent claims depend upon independent claim 1 and incorporate all the respective features of independent claim 1. Accordingly, the rejections of the dependent claims pursuant to 35 U.S.C. § 102 are deficient for at least the same respective rationale as applied to respective features of each of independent claim 1, and Applicant respectfully requests the rejections be withdrawn. Based on the above explanations of Bijalwan, it is respectfully submitted that Bijalwan does not describe the recited features of claims 11 and 20. Accordingly, it is respectfully submitted that claim 11, and claims 12-19 depending therefrom, are not anticipated by Bijalwan. Regarding the Applicant’s argument that the dependent claims are allowable at least due in part to their dependency on the independent claims, the Examiner respectfully disagrees and notes the instant rejections and response to arguments regarding the independent claims above. On page 16, Applicant argues: The rejected dependent claims respectively depend upon independent claims 1 and 11 and incorporate all the respective features of independent claims 1 and 11. Accordingly, the rejections of dependent claims 6, 10, 16, and 19 to 35 U.S.C. § 103 is deficient for at least the same respective rationale as applied to respective features of each of independent claims 1 and 10, and Applicant respectfully requests the rejections be withdrawn. Regarding the Applicant’s argument that the dependent claims are allowable at least due in part to their dependency on the independent claims, the Examiner respectfully disagrees and notes the instant rejections and response to arguments regarding the independent claims above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Moayed et al. (“Skipout: An adaptive Layer-Level Regularization Framework for Deep Neural Networks”) also discloses a method of skipping training on layers that do not add improvement to the model. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KAITLYN R LAU whose telephone number is (571)272-1429. The examiner can normally be reached Monday - Thursday: 8:00 am - 6:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.R.L./ Examiner, Art Unit 2148 /PAUL M KNIGHT/ Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Jul 20, 2023
Application Filed
Apr 30, 2026
Non-Final Rejection mailed — §101, §102, §103
Jun 05, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §101, §102, §103
Sep 23, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688298
FEATURE SELECTION FOR CYBERSECURITY THREAT DISPOSITION
4y 7m to grant Granted Jul 21, 2026
Patent 12602431
METHODS FOR PERFORMING INPUT-OUTPUT OPERATIONS IN A STORAGE SYSTEM USING ARTIFICIAL INTELLIGENCE AND DEVICES THEREOF
3y 10m to grant Granted Apr 14, 2026
Patent 12572828
METHOD FOR INDUSTRY TEXT INCREMENT AND ELECTRONIC DEVICE
4y 5m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
60%
Grant Probability
99%
With Interview (+66.7%)
3y 11m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 10 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month