Prosecution Insights
Last updated: October 02, 2026
Application No. 18/479,723

TEST-TIME ADAPTATION VIA SELF-DISTILLED REGULARIZATION

Non-Final OA §103§112
Filed
Oct 02, 2023
Priority
Nov 10, 2022 — provisional 63/424,315
Examiner
GHIMIRE, PRAYUKTA NMN
Art Unit
4100
Tech Center
4100
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
4 currently pending
Career history
2
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the application filed on 11/10/2022. Claims 1-28 are pending in the application and have been examined. Information Disclosure Statement The information disclosure statement (IDS) submitted on 03/25/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Such claim limitation(s) is/are: means for adding, means for training, means for adapting, means for classifying in claim 8, means for training, means for dividing in claim 9, and means for determining, means for constraining in claim 14. Because this/these claim limitation(s) is/are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are not being interpreted to cover only the corresponding structure, material, or acts described in the specification as performing the claimed function, and equivalents thereof. If applicant intends to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to remove the structure, materials, or acts that performs the claimed function; or (2) present a sufficient showing that the claim limitation(s) does/do not recite sufficient structure, materials, or acts to perform the claimed function. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 8-14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 8 recites the limitations “means for adding” and “means for classifying”. Claim 9 recites the limitation “means for dividing”. Both claims invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The disclosure is devoid of any means or step to perform the addition of the auxiliary network of a plurality of auxiliary networks to each partition of the plurality of partitions associated with a main network, classification of an input received at a model based on adapting each of the plurality of auxiliary network, and division of the main network into partitions. Paragraph 59 discloses that the encoder is segmented into K parts, which may be referred to as a model partition factor K where the auxiliary network is attached to each part K. However, it is not clearly explained how the auxiliary network is added. Paragraph 52 discloses convolutional layers, logistic regression layer, and a classification score (which may be a probability of the input data); however, it fails to clearly mention how the adapted auxiliary network is used in classification. As mentioned above, paragraph 59 discloses that the encoder is segmented into K parts, which may be referred to as a model partition factor K where the auxiliary network is attached to each part K. However, the specification also fails to clearly explain how the division is performed. Claims 10-14 are rejected as being dependent on rejected base claim 8 and claim 9 without curing any deficiencies. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 6-11, 13-18, 20-25, 27 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Niu et al. “Efficient Test-Time Model Adaptation without Forgetting”, in view of Choi (U.S PG Pub 2023/0306242). Regarding Claim 1, Niu discloses, A computer-implemented method, comprising: training each of the plurality of auxiliary networks with a training data to adapt to a test distribution; (Niu, section 3, “Without loss of generality, let P(x) be the distribution of training data x i ⅈ = 1 N   (namely x i ∼P(x))and f Θ ° ( x ) be a base model trained on labeled training data{( x i , y i ) } ⅈ = 1 N , where Θ ° denotes the model parameters.” and pg.3, figure 1 description, “Given a trained base model f Θ ° , we perform test-time adaptation with a model f Θ that initialized form Θ ° ”) adapting each of the plurality of auxiliary networks with test data to adapt to the test distribution; (Niu, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample.” and classifying an input received at a model based on adapting each of the plurality of auxiliary networks, the model including the main network and the plurality of auxiliary networks (Niu, section 5, “We conduct experiments on three benchmarks datasets for OOD generalization, i.e., CIFAR 10-C, ImageNet-C and ImageNet-R” specifying the input and Niu, page 12, section A.1, “ImageNet-R contains 30,000 images with various artistic renditions of 200 ImageNet classes, which are primarily collected from Flickr and filtered by Amazon MTurk annotators.” suggesting that the ImageNet dataset contains several classes.) However, Niu fails to disclose adding the auxiliary network of a plurality of auxiliary network to each partition of partitions associated with the main network. Choi, in the same field of endeavor as Niu (techniques for operating neural network models for improving model performance), teaches, adding a respective auxiliary network of a plurality of auxiliary networks to each partition of a plurality of partitions associated with a main network; (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers ) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” ) Niu and Choi are both considered to be analogous to the claimed invention because they are in the same field of operating neural network models for improving model performance. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train each of the auxiliary networks with training data and adapt each of the auxiliary network with test data to adapt to the test distribution while classifying the input received at the base model. The motivation to do wo would be to “tackle potential distribution shifts between training and test data” (Niu, Abstract). Regarding Claim 2, the Niu/Choi combination of claim 1 teaches, the computer-implemented method of claim 1, (and thus the rejection of claim 1 is incorporated). The combination via Choi further discloses, further comprising: training the main network with the training data; (Choi, Abstract, “A computing apparatus includes one or more processors, storage hardware storing instructions configured to, when executed by the one or more processors, cause the one or more processors to: extract calibration data from training data that is for training a main neural network, based on the calibration data” suggesting that the training data is used to train the main network.) and dividing the main network into the plurality of partitions (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” suggesting the presence of multiple layers of the main network) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train the main network with the training data and divide the main network into plurality of partitions. The motivation to divide the main network to have layers is because “layers of the neural network have different respective statistical characteristics of input data” (Choi, paragraph 4) Regarding Claim 3, the Niu/Choi combination of claim 2 teaches, the computer-implemented method of claim 2, (and thus the rejection of claim 2 is incorporated). The combination via Niu further discloses, wherein the main network is fixed after training the training data (Niu, pg. 3, figure 1 description, “During the adaptation process, we only update the parameters of batch normalization layers in f Θ and froze the rest parameters” suggesting that the main network is frozen/fixed.) Regarding Claim 4, the Niu/Choi combination of claim 1 teaches, the computer-implemented method of claim 1, (and thus the rejection of claim 1 is incorporated). The combination via Niu further teaches, wherein each of the plurality of auxiliary networks includes a first batch normalization layer and a convolution block. (Niu, pg.3, figure 1, [AltContent: ][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale [AltContent: textbox (Convolution block)] [AltContent: textbox (Batch normalization layer)] figure shows the batch normalization layer and a convolution block) Regarding Claim 6, the Niu/Choi combination of claim 1 teaches, The computer-implemented method of claim 1 (and thus the rejection of claim 1 is incorporated). The combination via Niu further teaches, wherein the training data is different than the test data (Niu, pg.1, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample” suggesting that the testing data and training data are different.) Regarding Claim 7, the Niu/Choi combination of claim 1 teaches, The computer-implemented method of claim 1, (and thus the rejection of claim 1 is incorporated). The combination via Choi further teaches, further comprising: determining, for each auxiliary network of the plurality of auxiliary networks, a mean absolute error between a first output of the respective partition and a second output of the auxiliary network (Choi, paragraph 61, “ The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.”) and constraining the adapting of each auxiliary network of the plurality of auxiliary networks based on the mean absolute error (Choi, paragraph 61, “The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.” indicating that the auxiliary network is fine-tuned/adapted. ) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to determine mean absolute error between the output of partition and the output of auxiliary network and adapt the auxiliary network based on the mean absolute error. The motivation to do so would be to have better performance. “The example 400 of FIG. 4 indicates that a fine-tuned LUT may have better performance than a baseline LUT which is not fine-tuned.” (Choi, paragraph 87) Regarding Claim 8, Niu discloses, An apparatus, comprising: means for training each of the plurality of auxiliary networks with a training data to adapt to a test distribution; (Niu, section 3, “Without loss of generality, let P(x) be the distribution of training data x i ⅈ = 1 N   (namely x i ∼P(x))and f Θ ° ( x ) be a base model trained on labeled training data{( x i , y i ) } ⅈ = 1 N , where Θ ° denotes the model parameters.” and pg.3, figure 1 description, “Given a trained base model f Θ ° , we perform test-time adaptation with a model f Θ that initialized form Θ ° ”) means for adapting each of the plurality of auxiliary networks with test data to adapt to the test distribution; (Niu, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample.” and means for classifying an input received at a model based on adapting each of the plurality of auxiliary networks, the model including the main network and the plurality of auxiliary networks (Niu, section 5, “We conduct experiments on three benchmarks datasets for OOD generalization, i.e., CIFAR 10-C, ImageNet-C and ImageNet-R” specifying the input and Niu, page 12, section A.1, “ImageNet-R contains 30,000 images with various artistic renditions of 200 ImageNet classes, which are primarily collected from Flickr and filtered by Amazon MTurk annotators.” suggesting that the ImageNet dataset contains several classes) However, Niu fails to disclose adding the auxiliary network of a plurality of auxiliary network to each partition of partitions associated with the main network. Choi, in the same field of endeavor as Niu (techniques for operating neural network models for improving model performance), teaches, means for adding a respective auxiliary network of a plurality of auxiliary networks to each partition of a plurality of partitions associated with a main network; (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers ) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” ) Niu and Choi are both considered to be analogous to the claimed invention because they are in the same field of operating neural network models for improving model performance. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train each of the auxiliary networks with training data and adapt each of the auxiliary network with test data to adapt to the test distribution while classifying the input received at the base model. The motivation to do wo would be to “tackle potential distribution shifts between training and test data” (Niu, Abstract). Regarding Claim 9, the Niu/Choi combination of claim 8 teaches, the apparatus of claim 8, (and thus the rejection of claim 8 is incorporated). The combination via Choi further discloses, further comprising: means for training the main network with the training data; (Choi, Abstract, “A computing apparatus includes one or more processors, storage hardware storing instructions configured to, when executed by the one or more processors, cause the one or more processors to: extract calibration data from training data that is for training a main neural network, based on the calibration data” suggesting that the training data is used to train the main network.) and means for dividing the main network into the plurality of partitions (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” suggesting the presence of multiple layers of the main network) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train the main network with the training data and divide the main network into plurality of partitions. The motivation to divide the main network to have layers is because “layers of the neural network have different respective statistical characteristics of input data” (Choi, paragraph 4) Regarding Claim 10, the Niu/Choi combination of claim 9, teaches, the apparatus of claim 9, (and thus the rejection of claim 9 is incorporated). The combination via Niu further discloses, wherein the main network is fixed after training the training data (Niu, pg. 3, figure 1 description, “During the adaptation process, we only update the parameters of batch normalization layers in f Θ and froze the rest parameters” suggesting that the main network is frozen/fixed.) Regarding Claim 11, the Niu/Choi combination of claim 8 teaches, the apparatus of claim 8, (and thus the rejection of claim 8 is incorporated). The combination via Niu further teaches, wherein each of the plurality of auxiliary networks includes a first batch normalization layer and a convolution block. (Niu, pg.3, figure 1, [AltContent: ][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale [AltContent: textbox (Convolution block)] [AltContent: textbox (Batch normalization layer)] figure shows the batch normalization layer and a convolution block) Regarding Claim 13, the Niu/Choi combination of claim 8 teaches, The apparatus of claim 8 (and thus the rejection of claim 8 is incorporated). The combination via Niu further teaches, wherein the training data is different than the test data (Niu, pg.1, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample” suggesting that the testing data and training data are different.) Regarding Claim 14, the Niu/Choi combination of claim 8 teaches, The apparatus of claim 8, (and thus the rejection of claim 8 is incorporated). The combination via Choi further teaches, further comprising: means for determining, for each auxiliary network of the plurality of auxiliary networks, a mean absolute error between a first output of the respective partition and a second output of the auxiliary network (Choi, paragraph 61, “ The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.”) and means for constraining the adapting of each auxiliary network of the plurality of auxiliary networks based on the mean absolute error (Choi, paragraph 61, “The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.” indicating that the auxiliary network is fine-tuned/adapted. ) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to determine mean absolute error between the output of partition and the output of auxiliary network and adapt the auxiliary network based on the mean absolute error. The motivation to do so would be to have better performance. “The example 400 of FIG. 4 indicates that a fine-tuned LUT may have better performance than a baseline LUT which is not fine-tuned.” (Choi, paragraph 87) Regarding Claim 15, Niu discloses, train each of the plurality of auxiliary networks with a training data to adapt to a test distribution; (Niu, section 3, “Without loss of generality, let P(x) be the distribution of training data x i ⅈ = 1 N   (namely x i ∼P(x))and f Θ ° ( x ) be a base model trained on labeled training data{( x i , y i ) } ⅈ = 1 N , where Θ ° denotes the model parameters.” and pg.3, figure 1 description, “Given a trained base model f Θ ° , we perform test-time adaptation with a model f Θ that initialized form Θ ° ”) adapt each of the plurality of auxiliary networks with test data to adapt to the test distribution; (Niu, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample.” and classify an input received at a model based on adapting each of the plurality of auxiliary networks, the model including the main network and the plurality of auxiliary networks (Niu, section 5, “We conduct experiments on three benchmarks datasets for OOD generalization, i.e., CIFAR 10-C, ImageNet-C and ImageNet-R” specifying the input and Niu, page 12, section A.1, “ImageNet-R contains 30,000 images with various artistic renditions of 200 ImageNet classes, which are primarily collected from Flickr and filtered by Amazon MTurk annotators.” suggesting that the ImageNet dataset contains several classes.) However, Niu fails to disclose a processor and memories coupled with processor and storing operations when executed by the processor, causes the apparatus to add the auxiliary network of a plurality of auxiliary network to each partition of partitions associated with the main network. Choi, in the same field of endeavor as Niu (techniques for operating neural network models for improving model performance), teaches, An apparatus, comprising: one or more processors (Choi, paragraph 6, “In one general aspect, a computing apparatus includes: one or more processors;”) one or more memories coupled with the one or more processors and storing instructions operable, when executed by one or more processors, to cause the apparatus to: (Choi, paragraph 108, “In one example, a processor or computer includes, or is connected to, one or more memories storing instructions or software that are executed by the processor or computer.”) add a respective auxiliary network of a plurality of auxiliary networks to each partition of a plurality of partitions associated with a main network; (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers ) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” ) Niu and Choi are both considered to be analogous to the claimed invention because they are in the same field of operating neural network models for improving model performance. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train each of the auxiliary networks with training data and adapt each of the auxiliary network with test data to adapt to the test distribution while classifying the input received at the base model. The motivation to do wo would be to “tackle potential distribution shifts between training and test data” (Niu, Abstract). Regarding Claim 16, the Niu/Choi combination of claim 15, teaches, the apparatus of claim 15, wherein execution of the instructions (and thus the rejection of claim 15 is incorporated). The combination via Choi further discloses, further causes the apparatus to train the main network with the training data; (Choi, Abstract, “A computing apparatus includes one or more processors, storage hardware storing instructions configured to, when executed by the one or more processors, cause the one or more processors to: extract calibration data from training data that is for training a main neural network, based on the calibration data” suggesting that the training data is used to train the main network.) and divide the main network into the plurality of partitions (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” suggesting the presence of multiple layers of the main network) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train the main network with the training data and divide the main network into plurality of partitions. The motivation to divide the main network to have layers is because “layers of the neural network have different respective statistical characteristics of input data” (Choi, paragraph 4) Regarding Claim 17, the Niu/Choi combination of claim 16, teaches, the apparatus of claim 16, (and thus the rejection of claim 16 is incorporated). The combination via Niu further discloses, wherein the main network is fixed after training the training data (Niu, pg. 3, figure 1 description, “During the adaptation process, we only update the parameters of batch normalization layers in f Θ and froze the rest parameters” suggesting that the main network is frozen/fixed.) Regarding Claim 18, the Niu/Choi combination of claim 15, teaches, the apparatus of claim 15, (and thus the rejection of claim 15 is incorporated). The combination via Niu further teaches, wherein each of the plurality of auxiliary networks includes a first batch normalization layer and a convolution block. (Niu, pg.3, figure 1, [AltContent: ][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale [AltContent: textbox (Convolution block)] [AltContent: textbox (Batch normalization layer)] figure shows the batch normalization layer and a convolution block) Regarding Claim 20, the Niu/Choi combination of claim 15, teaches, The apparatus of claim 15 (and thus the rejection of claim 15 is incorporated). The combination via Niu further teaches, wherein the training data is different than the test data (Niu, pg.1, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample” suggesting that the testing data and training data are different.) Regarding Claim 21, the Niu/Choi combination of claim 15, teaches, The apparatus of claim 15, wherein execution of the instructions further cause apparatus to: (and thus the rejection of claim 15 is incorporated). The combination via Choi further teaches, determine, for each auxiliary network of the plurality of auxiliary networks, a mean absolute error between a first output of the respective partition and a second output of the auxiliary network (Choi, paragraph 61, “ The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.”) and constrain the adapting of each auxiliary network of the plurality of auxiliary networks based on the mean absolute error (Choi, paragraph 61, “The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.” indicating that the auxiliary network is fine-tuned/adapted. ) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to determine mean absolute error between the output of partition and the output of auxiliary network and adapt the auxiliary network based on the mean absolute error. The motivation to do so would be to have better performance. “The example 400 of FIG. 4 indicates that a fine-tuned LUT may have better performance than a baseline LUT which is not fine-tuned.” (Choi, paragraph 87) Regarding Claim 22, Niu discloses, program code to train each of the plurality of auxiliary networks with a training data to adapt to a test distribution; (Niu, section 3, “Without loss of generality, let P(x) be the distribution of training data x i ⅈ = 1 N   (namely x i ∼P(x))and f Θ ° ( x ) be a base model trained on labeled training data{( x i , y i ) } ⅈ = 1 N , where Θ ° denotes the model parameters.” and pg.3, figure 1 description, “Given a trained base model f Θ ° , we perform test-time adaptation with a model f Θ that initialized form Θ ° ”) program code to adapt each of the plurality of auxiliary networks with test data to adapt to the test distribution; (Niu, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample.” and program code to classify an input received at a model based on adapting each of the plurality of auxiliary networks, the model including the main network and the plurality of auxiliary networks (Niu, section 5, “We conduct experiments on three benchmarks datasets for OOD generalization, i.e., CIFAR 10-C, ImageNet-C and ImageNet-R” specifying the input and Niu, page 12, section A.1, “ImageNet-R contains 30,000 images with various artistic renditions of 200 ImageNet classes, which are primarily collected from Flickr and filtered by Amazon MTurk annotators.” suggesting that the ImageNet dataset contains several classes.) However, Niu fails to disclose a non-transitory computer-readable medium having program code executed by the processor, and the program code to add the auxiliary network of a plurality of auxiliary network to each partition of partitions associated with the main network. Choi, in the same field of endeavor as Niu (techniques for operating neural network models for improving model performance), teaches, A non-transitory computer-readable medium having program code recorded thereon, the program code executed by one or more processors (Choi, paragraph 108, “In one example, a processor or computer includes, or is connected to, one or more memories storing instructions or software that are executed by the processor or computer.”) program code to add a respective auxiliary network of a plurality of auxiliary networks to each partition of a plurality of partitions associated with a main network; (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers ) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” ) Niu and Choi are both considered to be analogous to the claimed invention because they are in the same field of operating neural network models for improving model performance. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train each of the auxiliary networks with training data and adapt each of the auxiliary network with test data to adapt to the test distribution while classifying the input received at the base model. The motivation to do wo would be to “tackle potential distribution shifts between training and test data” (Niu, Abstract). Regarding Claim 23, the Niu/Choi combination of claim 22, teaches, the non-transitory computer readable medium of claim 22, (and thus the rejection of claim 22 is incorporated). The combination via Choi further discloses, further comprises: program code to train the main network with the training data; (Choi, Abstract, “A computing apparatus includes one or more processors, storage hardware storing instructions configured to, when executed by the one or more processors, cause the one or more processors to: extract calibration data from training data that is for training a main neural network, based on the calibration data” suggesting that the training data is used to train the main network.) and program code to divide the main network into the plurality of partitions (Choi, paragraph 45, “In some embodiments, there may be auxiliary networks for respective layers of the main neural network (for some or all main layers) which may generate respective LUTs to serve as substitutes for non-linear functions of the respective layers of the main neural network.” suggesting the presence of multiple layers of the main network) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to train the main network with the training data and divide the main network into plurality of partitions. The motivation to divide the main network to have layers is because “layers of the neural network have different respective statistical characteristics of input data” (Choi, paragraph 4) Regarding Claim 24, the Niu/Choi combination of claim 23 teaches, the non-transitory computer-readable medium of claim 23, (and thus the rejection of claim 23 is incorporated). The combination via Niu further discloses, wherein the main network is fixed after training the training data (Niu, pg. 3, figure 1 description, “During the adaptation process, we only update the parameters of batch normalization layers in f Θ and froze the rest parameters” suggesting that the main network is frozen/fixed.) Regarding Claim 25, the Niu/Choi combination of claim 22, teaches, the non-transitory computer-readable medium of claim 22, (and thus the rejection of claim 22 is incorporated). The combination via Niu further teaches, wherein each of the plurality of auxiliary networks includes a first batch normalization layer and a convolution block. (Niu, pg.3, figure 1, [AltContent: ][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale [AltContent: textbox (Convolution block)] [AltContent: textbox (Batch normalization layer)] figure shows the batch normalization layer and a convolution block) Regarding Claim 27, the Niu/Choi combination of claim 22, teaches, the non-transitory computer-readable medium of claim 22, (and thus the rejection of claim 22 is incorporated). The combination via Niu further teaches, wherein the training data is different than the test data (Niu, pg.1, abstract, “Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample” suggesting that the testing data and training data are different.) Regarding Claim 28, the Niu/Choi combination of claim 22, teaches, the non-transitory computer-readable medium of claim 22, (and thus the rejection of claim 22 is incorporated). The combination via Choi further teaches, wherein the program code further comprises: program code to determine, for each auxiliary network of the plurality of auxiliary networks, a mean absolute error between a first output of the respective partition and a second output of the auxiliary network (Choi, paragraph 61, “ The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.”) and program code to constrain the adapting of each auxiliary network of the plurality of auxiliary networks based on the mean absolute error (Choi, paragraph 61, “The processor 200 may fine-tune the parameter of the auxiliary network based on a mean absolute error (MAE) between the output of the auxiliary network and the output of the non-linear function, for example.” indicating that the auxiliary network is fine-tuned/adapted. ) It would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu and Choi to determine mean absolute error between the output of partition and the output of auxiliary network and adapt the auxiliary network based on the mean absolute error. The motivation to do so would be to have better performance. “The example 400 of FIG. 4 indicates that a fine-tuned LUT may have better performance than a baseline LUT which is not fine-tuned.” (Choi, paragraph 87) Claims 5,12,19 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Niu, in view of Choi, and further in view of Park et al. “A review and comparison of convolution neural network models under a unified framework” Regarding Claim 5, the Niu/Choi combination of claim 4 teaches, The computer-implemented method of claim 4, (and thus the rejection of claim 4 is incorporated). The combination via Niu further teaches, wherein the convolution block includes a second batch normalization layer (pg.3, figure 1, [AltContent: rect][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale ) [AltContent: textbox (Multiple batch normalization layers)] However, Choi/Niu combination fails to mention the convolution layer, and a ReLU. Park, who is also in the same field of endeavor as Choi and Niu (techniques relating to evaluation and optimization of neural network models), teaches, a convolution layer, and a rectified linear unit (ReLU). (Park, pg.163, section 2.2, “In Figure 5, ‘Conv Block’ and ‘Id Block’ stand for the convolution block and identity block, respectively [AltContent: textbox (Convolution block)][AltContent: rect] PNG media_image2.png 152 505 media_image2.png Greyscale ” and section 2.3, “Zagoruyko and Komodakis (2016) proposed wide residual networks (WRN) based on the structure of ResNet (He et al., 2016)…. Moreover, WRN changed the order of convolution, batch normalization (BN), and activation (ReLU) into BN-ReLU-Conv from the typical order Conv-BN-ReLU.” showing that the convolution block included ReLU and Convolution layer) Niu, Choi and Park are all considered to be analogous to the claimed invention because they are in the same field of endeavor which relates to evaluation and optimization of neural network models. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu, Choi and Park to include the convolution layer, a second batch normalization layer, and a rectified linear unit (ReLU) in the convolution block. The reason for doing so would be to have faster training and have higher accuracy (Park, section 2.3). Regarding Claim 12, the Niu/Choi combination of claim 11, teaches, The apparatus of claim 11, (and thus the rejection of claim 11 is incorporated). The combination via Niu further teaches, wherein the convolution block includes a second batch normalization layer (pg.3, figure 1, [AltContent: rect][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale ) [AltContent: textbox (Multiple batch normalization layers)] However, Choi/Niu combination fails to mention the convolution layer, and a ReLU. Park, who is also in the same field of endeavor as Choi and Niu (techniques relating to evaluation and optimization of neural network models), teaches, a convolution layer, and a rectified linear unit (ReLU). (Park, pg.163, section 2.2, “In Figure 5, ‘Conv Block’ and ‘Id Block’ stand for the convolution block and identity block, respectively [AltContent: textbox (Convolution block)][AltContent: rect] PNG media_image2.png 152 505 media_image2.png Greyscale ” and section 2.3, “Zagoruyko and Komodakis (2016) proposed wide residual networks (WRN) based on the structure of ResNet (He et al., 2016)…. Moreover, WRN changed the order of convolution, batch normalization (BN), and activation (ReLU) into BN-ReLU-Conv from the typical order Conv-BN-ReLU.” showing that the convolution block included ReLU and Convolution layer) Niu, Choi and Park are all considered to be analogous to the claimed invention because they are in the same field of endeavor which relates to evaluation and optimization of neural network models. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu, Choi and Park to include the convolution layer, a second batch normalization layer, and a rectified linear unit (ReLU) in the convolution block. The reason for doing so would be to have faster training and have higher accuracy (Park, section 2.3). Regarding Claim 19, the Niu/Choi combination of claim 18, teaches, The apparatus of claim 18, (and thus the rejection of claim 18 is incorporated). The combination via Niu further teaches, wherein the convolution block includes a second batch normalization layer (pg.3, figure 1, [AltContent: rect][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale ) [AltContent: textbox (Multiple batch normalization layers)] However, Choi/Niu combination fails to mention the convolution layer, and a ReLU. Park, who is also in the same field of endeavor as Choi and Niu (techniques relating to evaluation and optimization of neural network models), teaches, a convolution layer, and a rectified linear unit (ReLU). (Park, pg.163, section 2.2, “In Figure 5, ‘Conv Block’ and ‘Id Block’ stand for the convolution block and identity block, respectively [AltContent: textbox (Convolution block)][AltContent: rect] PNG media_image2.png 152 505 media_image2.png Greyscale ” and section 2.3, “Zagoruyko and Komodakis (2016) proposed wide residual networks (WRN) based on the structure of ResNet (He et al., 2016)…. Moreover, WRN changed the order of convolution, batch normalization (BN), and activation (ReLU) into BN-ReLU-Conv from the typical order Conv-BN-ReLU.” showing that the convolution block included ReLU and Convolution layer) Niu, Choi and Park are all considered to be analogous to the claimed invention because they are in the same field of endeavor which relates to evaluation and optimization of neural network models. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu, Choi and Park to include the convolution layer, a second batch normalization layer, and a rectified linear unit (ReLU) in the convolution block. The reason for doing so would be to have faster training and have higher accuracy (Park, section 2.3). Regarding Claim 26, the Niu/Choi combination of claim 25, teaches, The non-transitory computer-readable medium of claim 25, (and thus the rejection of claim 25 is incorporated). The combination via Niu further teaches, wherein the convolution block includes a second batch normalization layer (pg.3, figure 1, [AltContent: rect][AltContent: rect] PNG media_image1.png 305 1023 media_image1.png Greyscale ) [AltContent: textbox (Multiple batch normalization layers)] However, Choi/Niu combination fails to mention the convolution layer, and a ReLU. Park, who is also in the same field of endeavor as Choi and Niu (techniques relating to evaluation and optimization of neural network models), teaches, a convolution layer, and a rectified linear unit (ReLU). (Park, pg.163, section 2.2, “In Figure 5, ‘Conv Block’ and ‘Id Block’ stand for the convolution block and identity block, respectively [AltContent: textbox (Convolution block)][AltContent: rect] PNG media_image2.png 152 505 media_image2.png Greyscale ” and section 2.3, “Zagoruyko and Komodakis (2016) proposed wide residual networks (WRN) based on the structure of ResNet (He et al., 2016)…. Moreover, WRN changed the order of convolution, batch normalization (BN), and activation (ReLU) into BN-ReLU-Conv from the typical order Conv-BN-ReLU.” showing that the convolution block included ReLU and Convolution layer) Niu, Choi and Park are all considered to be analogous to the claimed invention because they are in the same field of endeavor which relates to evaluation and optimization of neural network models. Therefore, it would’ve been obvious to someone of ordinary skill in art before the filling date of the claimed invention to have combined the teachings of Niu, Choi and Park to include the convolution layer, a second batch normalization layer, and a rectified linear unit (ReLU) in the convolution block. The reason for doing so would be to have faster training and have higher accuracy (Park, section 2.3). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PRAYUKTA GHIMIRE whose telephone number is (571)270-5484. The examiner can normally be reached M-F, 9 am to 5 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /P.N.G./Examiner, Art Unit 2122 /MICHAEL H HOANG/PRIMARY EXAMINER, Art Unit 2122
Read full office action

Prosecution Timeline

Oct 02, 2023
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month