Prosecution Insights
Last updated: October 02, 2026
Application No. 18/090,211

METHOD AND APPARATUS FOR CONSTRUCTING DOMAIN ADAPTIVE NETWORK

Final Rejection §103
Filed
Dec 28, 2022
Priority
Feb 14, 2022 — RE 10-2022-0018857
Examiner
DEVORE, CHRISTOPHER DILLON
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Electronics and Telecommunications Research Institute
OA Round
2 (Final)
57%
Grant Probability
Moderate
3-4
OA Rounds
5m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 57% of resolved cases
57%
Career Allowance Rate
12 granted / 21 resolved
+2.1% vs TC avg
Strong +36% interview lift
Without
With
+36.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
10 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
29.3%
-10.7% vs TC avg
§103
44.3%
+4.3% vs TC avg
§102
7.0%
-33.0% vs TC avg
§112
17.8%
-22.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 21 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Remarks Remarks page 8-12, Applicant contends: Amended claim limitations more clearly define statutory subject matter. Response: Applicant’s arguments regarding 101 are seen as convincing, as the claims appear directed towards an improvement of a computer by improving a machine learning model. As a result, the 101 rejections are removed. Remarks page 12-15, Applicant contends: 103 rejections are traversed by amendments and currently recited prior art not teaching every limitation. Response: The recited prior art is seen as reciting elements of the claimed limitations, as the limitations rolled up into claim 1 are still taught by relevant sections used in previous rejections. As recited by office action claim rejections, weight vector [Hashem 3. Linear Combinations of Neural Networks page 3 and related support in claim 1 rejection], input data from prototype domains [Samdani 3 Feature Ensembles for Domain Adaptation page 2 in claim 1 rejection], neural network pool [Samdani 3 Feature Ensembles for Domain Adaptation page 2 in claim 1 rejection], and linear combination using weight vector to create final neural network [Hashem Introduction page 1 and Izmailov 3.5 Connection To Ensembling page 7 in claim 1 rejection] are still considered taught by the currently recited art. The composite domain and continuous variation for the input data have not been previously examined. Applicant’s arguments with respect to claim(s) 1, 10, or 14 have been considered but are moot because the new ground of rejection contain elements that have not been previously examined (such as elements pertaining to continuous variation or composite domain) or does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3, 4, 7-10, 14, 16, 17, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hashem (“Optimal Linear Combinations of Neural Networks”), referred to as Hashem in this document, and further in combination with Izmailov et al (“Averaging Weights Leads to Wider Optima and Better Generalization”), referred to as Izmailov in this document, and further in combination with Samdani et al (“Domain Adaptation with Ensemble of Feature Groups”), referred to as Samdani in this document, and further in combination with Yeo et al (US 20220036152 A1), referred to as Yeo in this document. Regarding Claim 1: Hashem teaches: determine a weight vector to be applied to one or more neural networks stored in a neural network pool based on input data, [Hashem 3. Linear Combinations of Neural Networks page 3]: “Consider a multi-input single-output mapping approximated by a trained NN. A trained NN accepts a vector-valued input x and returns a scalar output (response) y(x) [determine a weight vector to be applied to one or more neural networks stored in a neural network pool based on input data]. The approximation error is δ(x) = r(x) − y(x), where r(x) is the response of the real system (true response) for x.” Support in Hashem indicating the weight can be a weight vector [Hashem 3. Linear Combinations of Neural Networks page 3]: “One approach for the multi-output case is to compute an optimal combination-weight vector for each output separately” construct a final neural network by performing linear combination of parameters of the one or more neural networks using the weight vector, [Hashem Introduction page 1]: “Hashem and Schmeiser (1995) proposed forming a linear combination of the corresponding outputs of the trained NNs, instead of just using the apparent best network. Combining the trained networks may help integrate the knowledge acquired by the component networks and often produces superior model accuracy compared to the single best-trained network (Hashem, 1993; Hashem and Schmeiser, 1993; Hashem et al. (1993), Hashem et al. (1994)). Optimal linear combinations (OLCs) of neural networks are constructed by forming weighted sums [by performing linear combination of parameters of the one or more neural networks using the weight vector] of the corresponding outputs of the networks.” and output result data of the input data using the final neural network, [Hashem 2 Related Work page 3]: “For a given input, x, the output [and output result data of the input data using the final neural network] of the combined model y, is the weighted sum of the corresponding outputs of the component NNs, yj, j = 1,...,p, and the aj’s are the associated combination-weights.” Hashem does not explicitly teach: An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data; and a processor configured to control the memory, wherein the processor is configured to stored in a neural network pool construct a final neural network (Hashem notes a combined model [Hashem 2 Related Work page 3] which appears to fit the limitation under BRI, but another reference is used in hopes of progressing prosecution faster) and the one or more neural networks are trained using data corresponding to each prototype domain the input data is data associated with one or more prototype domains and includes data associated with a composite domain formed by a combination of the one or more prototype domains or a domain representing a continuous variation of the one or more prototype domains Yeo teaches: An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data; and a processor configured to control the memory, wherein the processor is configured to [Yeo 0164]: “The loaded dedicated artificial intelligence model may be a compressed artificial intelligence model having a smaller size than the generic-purpose artificial intelligence model stored in the memory 110. As described above, the processor 120 may load a dedicated artificial intelligence model having a small size and may perform an operation, and thus an operation amount for the target data can be reduced, and the processing speed can be improved, and the resources (e.g., the memory [An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data], the CPU [and a processor configured to control the memory, wherein the processor is configured to], the GPU, etc.) of the electronic apparatus 100 are not wasted, and thus usefulness of the resources can be enhanced.” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Yeo. Hashem and Yeo are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Yeo in order to be able to create a physical embodiment of the invention or to give the invention physical form [Yeo 0164]: “(and the resources (e.g., the memory, the CPU, the GPU, etc.) of the electronic apparatus 100 are not wasted, and thus usefulness of the resources can be enhanced)”. Izmailov teaches: construct a final neural network [Izmailov 3.5 Connection To Ensembling page 7]: “In SWA instead of averaging the predictions of the models we average their weights [construct a final neural network]. However, the predictions proposed by FGE ensembles and SWA models have similar properties.” Support for interpretation of combining weights/parameters of models [Current Application page 4 line 18]: “In addition, the final neural network may be derived based on a linear combination of parameters of the one or more neural networks using the weight.”. Izmailov notes the methods of ensembling outputs/predictions and the weights are similar, thus supporting the motivation of the combination. One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Izmailov. Hashem and Izmailov are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Izmailov in order to be able to create a better performing model by combining weights instead of predictions ([Izmailov 4.1 Cifar Datasets page 9]: “Amazingly, SWA is able to achieve comparable or better performance than FGE ensembles with just one model.”). Samdani teaches: stored in a neural network pool [Samdani 3 Feature Ensembles for Domain Adaptation page 2]: “Compared to existing domain adaption methods, FEAD provides an additional degree of freedom to adjust the trade-off between the generalization error and domain distribution change of individual classifiers [stored in a neural network pool, as a network pool is interpreted as a way to refer to multiple or a collection of neural networks as shown by Figure 3 of the current application] trained on the corresponding feature groups.” and the one or more neural networks are trained using data corresponding to each prototype domain [Samdani 3 Feature Ensembles for Domain Adaptation page 2]: “Compared to existing domain adaption methods, FEAD provides an additional degree of freedom to adjust the trade-off between the generalization error and domain distribution change of individual classifiers trained on the corresponding feature groups [and the one or more neural networks are trained using data corresponding to each prototype domain as Samdani teaches aspects for domain adaptation and notes within this quote that classifiers can be trained on information for a form of domain, such a feature group.].” Samdani is relevant for the combination of models as shown by [Samdani Introduction page 1]: “Given a set of feature groups that capture this notion, where the grouping can be decided by domain knowledge or statistics derived from unlabeled data, we first train individual classifiers separately using only the corresponding group of features. The final model is a weighted ensemble of individual classifiers, where the weights are tuned based on the performance of the ensemble on a small amount of labeled target data.”. the input data is data associated with one or more prototype domains [Samdani 3 Feature Ensembles for Domain Adaptation page 2]: “Compared to existing domain adaption methods, FEAD provides an additional degree of freedom to adjust the trade-off between the generalization error and domain distribution change of individual classifiers trained on the corresponding feature groups [he input data is data associated with one or more prototype domains as Samdani teaches aspects for domain adaptation and notes within this quote that classifiers can be trained on information for a form of domain, such as a feature group.].” and includes data associated with a composite domain formed by a combination of the one or more prototype domains [Samdani Introduction page 1]: “As observed by previous studies, unfortunately, na¨ıvely applying classifiers to a different domain often leads to considerable performance degradation [Daum´e III, 2007; Jiang and Zhai, 2007]. Consequently, a number of approaches have been proposed recently to address the problem of domain adaptation. Some methods focus on re-weighting training instances from different domains [and includes data associated with a composite domain formed by a combination of the one or more prototype domains] [Jiang and Zhai, 2007; Bickel et al., 2009], while others modify the feature space to capture domain-specific and domain-invariant aspects [Blitzer et al., 2006; Daum´e III, 2007; Finkel and Manning, 2009; Jiang and Zhai, 2006]. Despite the fact that these methods adapt seemingly very different strategies, the shared rationale behind is to bring the empirical source distribution closer to the target domain, and thus increase the accuracy of the classifier when evaluated on the target domain data.” Composite domain is interpreted as a complex domain, as the specification does not describe a composite domain but the description of composite domain in the claim matches that of the complex domain given in the specification. ([Current Specification page 3 line 2]: "For example, assuming that there is a neural network trained with data collected in a dark night environment and a neural network trained with data collected in a bright day environment, it becomes difficult to determine which neural network should be used in the dusky evening when the sun goes down. In addition, even if the neural network learns each domain, such as a rainy environment and an evening environment, the same problem also occurs when a complex situation such as a rainy evening occurs. To this end, when the neural network for the evening environment is additionally trained in addition to the neural network for the rainy environment, a problem arises that increases the burden of time and resources again.") or a domain representing a continuous variation of the one or more prototype domains [Samdani 4.2 page 5]: “Note that we are making a distinction between source and target data by cutting off a continuous timeline [or a domain representing a continuous variation of the one or more prototype domains] at an arbitrarily picked point. We assume that this distinction is a good enough approximation to the ‘source-target’ setting of domain adaption.” Some support for the above limitations having domains that can different than another is also indicated by Samdani noting cross-domain aspects ([Samdani Introduction page 1]: “Given a set of feature groups that capture this notion, where the grouping can be decided by domain knowledge or statistics derived from unlabeled data, we first train individual classifiers separately using only the corresponding group of features. The final model is a weighted ensemble of individual classifiers, where the weights are tuned based on the performance of the ensemble on a small amount of labeled target data. Compared to the existing approaches, our method is unique in that it considers the cross-domain behavior of different feature groups directly in terms of their classification accuracy. Afterwards, instead of creating an instance distribution close to the target domain, it adjusts the influence of features on the final classifier by tuning their weights.”) Continuous variation is interpreted as referring to consecutive domains, as the specification does not describe continuous variation, but the idea of a continuous variation appears to match the premise of consecutive domains. ([Current Specification page 20 line 17]: "More specifically, FIG. 6 is one of the embodiments of a method of securing data for learning consecutive domains, that is, one or more domains and a learning method of a combiner... As an example, as mentioned above, when the neural network is to be trained as a neural network that performs image-related processing, a change 702 may be made to information included in the image, that is, environmental elements such as rainfall, an amount of sunlight, and the degree of fog. That is, synthetic data may be generated by changing the weather, time zone, etc., of the image. Each of these environmental elements may be treated as a separate prototyping domain, respectively.") One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Samdani. Hashem and Samdani are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Samdani in order to be able to utilize domain adaptation to create more relevant or accurate models ([Samdani Introduction page 1]: “However, in several real-world applications, it is often highly desirable to train a classifier from one source domain, and apply it to a similar but different target domain, where the annotated data is unavailable or expensive to create. One example of this scenario is to learn a text categorizer from a large collection of labeled newswire articles, but use it to process regular Web documents. In this domain adaptation setting, the goal is to leverage the data available in the source domain to improve the accuracy of the model when testing on target domain examples.”) and ([Samdani Introduction page 1]: “Despite the fact that these methods adapt seemingly very different strategies, the shared rationale behind is to bring the empirical source distribution closer to the target domain, and thus increase the accuracy of the classifier when evaluated on the target domain data.”) Regarding Claim 3: The apparatus of claim 1 is taught by Hashem, Izmailov, Samdani, and Yeo. Hashem teaches: wherein the one or more neural networks all have the same structure [Hashem 2. Related Work page 2]: “Hansen and Salamon (1990) suggested training a group of networks of the same architecture [wherein the one or more neural networks all have the same structure] but initialized with different connection-weights. Then, as screen subset of the trained networks is used for making the final classification decision by some voting scheme.” Regarding Claim 4: The apparatus of claim 1 is taught by Hashem, Izmailov, Samdani, and Yeo. Yeo teaches: the neural network pool is compressed through a singular vector decomposition (SVD) technique [Yeo 0013]: “FIG. 3A is a diagram showing compression of an artificial intelligence model using an SVD algorithm [and the neural network pool is compressed through a singular vector decomposition (SVD) technique] according to an embodiment” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Yeo. Hashem and Yeo are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Yeo in order to be able to compress neural networks to preserve resources or use resources more efficiently ([Yeo 0164]: “The loaded dedicated artificial intelligence model may be a compressed artificial intelligence model having a smaller size than the generic-purpose artificial intelligence model stored in the memory 110. As described above, the processor 120 may load a dedicated artificial intelligence model having a small size and may perform an operation, and thus an operation amount for the target data can be reduced, and the processing speed can be improved, and the resources (e.g., the memory, the CPU, the GPU, etc.) of the electronic apparatus 100 are not wasted, and thus usefulness of the resources can be enhanced.”) Regarding Claim 7: The apparatus of claim 1 is taught by Hashem, Izmailov, Samdani, and Yeo. Izmailov teaches: wherein the one or more neural networks are derived based on a primitive neural network trained using each prototype domain [Izmailov 3.2 SWA Algorithm]: “Following Garipov et al. [2018], we start with a pretrained model wˆ [wherein the one or more neural networks are derived based on a primitive neural network trained using each prototype domain]. We will refer to the number of epochs required to train a given DNN with the conventional training procedure as its training budget and will denote it by B. The pretrained model wˆ can be trained with the conventional training procedure for full training budget or reduced number of epochs (e.g. 0.75B). In the latter case we just stop the training early without modifying the learning rate schedule. Starting from wˆ we continue training, using a cyclical or constant learning rate schedule. When using a cyclical learning rate we capture the models wi that correspond to the minimum values of the learning rate (see Figure 2), following Garipov et al. [2018]. For constant learning rates we capture models at each epoch. Next, we average the weights of all the captured networks wi to get our final model wSWA.” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Izmailov for the same motivation provided in claim 1 to combine with Izmailov, as well as the derivation from a primitive neural network acts as a starting point for method of providing a better model, thus still sensible for one of ordinary skill to combine. Regarding Claim 8: The apparatus of claim 7 is taught by Hashem, Izmailov, Samdani, and Yeo. Hashem teaches: wherein the primitive neural network is trained through supervised learning or representation learning [Hashem Introduction page 2]: “This paper focuses mainly on function approximation or regression problems. However, MSE-OLCs are also applicable to supervised [wherein the primitive neural network is trained through supervised learning or representation learning as Hashem is teaching in this quote supervised learning in the example of supervised classification, where the primitive neural network is taught in claim 7] classification problems” Regarding Claim 9: The apparatus of claim 1 is taught by Hashem, Izmailov, Samdani, and Yeo. Hashem teaches: wherein the weight vector is derived based on a multilayer neural network, [Hashem Introduction page 2]: “The class of neural networks investigated here is the class of multilayer [wherein the weight vector is derived based on a multilayer neural network as this quote teaches multilayer neural networks as claim 1 teaches the determination of a weight] feedforward networks. No further assumptions regarding the network architecture or the learning method are needed” and the multilayer neural network is trained based on a weighted sum of results of the one or more neural networks [Hashem 3. Linear Combinations of Neural Networks page 3]: “Consider a multi-input single-output mapping approximated by a trained NN. A trained NN accepts a vector-valued input x and returns a scalar output (response) y(x). The approximation error is δ(x) = r(x) − y(x), where r(x) is the response of the real system (true response) for x [and the multilayer neural network is trained based on a weighted sum of results of the one or more neural networks as Hashem is teaching that the network for predicting a weight is trained using the output of one or more neural networks].” Further support is given in the words following the quote [Hashem 3. Linear Combinations of Neural Networks page 3]: “According to (Hashem& Schmeiser,1995),a linear combination of the outputs of p NNs returns the scalar y(x; a)= ∑ j p a j y j ( x ) , with the corresponding approximation error (x; a)= r(x) –y(x; a); where yj(x) is the output of the jth network and aj is the associated combination-weight” Regarding Claim 10: Hashem teaches: perform multilayer neural network learning to determine a weight vector to be applied to one or more neural networks using the collected learning data, and the weight is derived to combine result values of the one or more neural networks [Hashem 3. Linear Combinations of Neural Networks page 3]: “Consider a multi-input single-output mapping approximated by a trained NN. A trained NN accepts a vector-valued input x and returns a scalar output (response) y(x) [perform multilayer neural network learning to determine a weight vector to be applied to one or more neural networks using the collected learning data,]. The approximation error is δ(x) = r(x) − y(x), where r(x) is the response of the real system (true response) for x.” Further support is given in the words following the quote [Hashem 3. Linear Combinations of Neural Networks page 3]: “According to (Hashem& Schmeiser,1995), a linear combination of the outputs of p NNs returns the scalar y(x; a)= ∑ j p a j y j ( x ) , with the corresponding approximation error (x; a)= r(x) –y(x; a); where yj(x) is the output of the jth network and aj is the associated combination-weight” Support in Hashem indicating the weight can be a weight vector [Hashem 3. Linear Combinations of Neural Networks page 3]: “One approach for the multi-output case is to compute an optimal combination-weight vector for each output separately” construct a final neural network by performing linear combination of parameters of the one or more neural networks using the weight vector [Hashem Introduction page 1]: “Hashem and Schmeiser (1995) proposed forming a linear combination of the corresponding outputs of the trained NNs, instead of just using the apparent best network. Combining the trained networks may help integrate the knowledge acquired by the component networks and often produces superior model accuracy compared to the single best-trained network (Hashem, 1993; Hashem and Schmeiser, 1993; Hashem et al. (1993), Hashem et al. (1994)). Optimal linear combinations (OLCs) of neural networks are constructed by forming weighted sums [by performing linear combination of parameters of the one or more neural networks using the weight vector] of the corresponding outputs of the networks.” Hashem does not explicitly teach: An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data; and a processor configured to control the memory, wherein the processor is configured to collect learning data associated with one or more prototype domains, construct a final neural network (Hashem notes a combined model [Hashem 2 Related Work page 3] which appears to fit the limitation under BRI, but another reference is used in hopes of progressing prosecution faster) the learning data includes data associated with a composite domain formed by a combination of the one or more prototype domains or a domain representing a continuous variation of the one or more prototype domains Yeo teaches: An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data; and a processor configured to control the memory, wherein the processor is configured to [Yeo 0164]: “The loaded dedicated artificial intelligence model may be a compressed artificial intelligence model having a smaller size than the generic-purpose artificial intelligence model stored in the memory 110. As described above, the processor 120 may load a dedicated artificial intelligence model having a small size and may perform an operation, and thus an operation amount for the target data can be reduced, and the processing speed can be improved, and the resources (e.g., the memory [An apparatus for constructing a domain adaptive network, the apparatus comprising: a memory configured to store data;], the CPU [and a processor configured to control the memory, wherein the processor is configured to], the GPU, etc.) of the electronic apparatus 100 are not wasted, and thus usefulness of the resources can be enhanced.” The motivation to combine with Yeo is the same motivation to combine with Yeo as claim 1. Samdani teaches: collect learning data associated with one or more prototype domains, [Samdani 3 Feature Ensembles for Domain Adaptation page 2]: “Compared to existing domain adaption methods, FEAD provides an additional degree of freedom to adjust the trade-off between the generalization error and domain distribution change of individual classifiers trained on the corresponding feature groups [collect learning data associated with one or more prototype domains as Samdani teaches aspects for domain adaptation and notes within this quote that classifiers can be trained on information for a form of domain, such a feature group.].” Where further support for collecting data is shown by Samdani noting data is collected for datasets [Samdani 3.3 FEAD as Product of Experts page 4]: “In this set of experiments, we use the benchmark dataset released by Blitzer et al. [2007], which consists of reviews of four different product categories: books, DVDs, electronics, and kitchen appliances, collected from Amazon.com.” the learning data includes data associated with a composite domain formed by a combination of the one or more prototype domains [Samdani Introduction page 1]: “As observed by previous studies, unfortunately, na¨ıvely applying classifiers to a different domain often leads to considerable performance degradation [Daum´e III, 2007; Jiang and Zhai, 2007]. Consequently, a number of approaches have been proposed recently to address the problem of domain adaptation. Some methods focus on re-weighting training instances from different domains [the learning data includes data associated with a composite domain formed by a combination of the one or more prototype domains] [Jiang and Zhai, 2007; Bickel et al., 2009], while others modify the feature space to capture domain-specific and domain-invariant aspects [Blitzer et al., 2006; Daum´e III, 2007; Finkel and Manning, 2009; Jiang and Zhai, 2006]. Despite the fact that these methods adapt seemingly very different strategies, the shared rationale behind is to bring the empirical source distribution closer to the target domain, and thus increase the accuracy of the classifier when evaluated on the target domain data.” or a domain representing a continuous variation of the one or more prototype domains [Samdani 4.2 page 5]: “Note that we are making a distinction between source and target data by cutting off a continuous timeline [or a domain representing a continuous variation of the one or more prototype domains] at an arbitrarily picked point. We assume that this distinction is a good enough approximation to the ‘source-target’ setting of domain adaption.” The motivation to combine with Samdani is the same motivation as the motivation to combine with Samdani in claim 1. Izmailov teaches: construct a final neural network [Izmailov 3.5 Connection To Ensembling page 7]: “In SWA instead of averaging the predictions of the models we average their weights [construct a final neural network]. However, the predictions proposed by FGE ensembles and SWA models have similar properties.” Support for interpretation of combining weights/parameters of models [Current Application page 4 line 18]: “In addition, the final neural network may be derived based on a linear combination of parameters of the one or more neural networks using the weight.”. Izmailov notes the methods of ensembling outputs/predictions and the weights are similar, thus supporting the motivation of the combination. The motivation to combine with Izmailov is the same motivation as to combine with Izmailov in claim 1. Regarding Claim 14: Claim 14 is analogous to claim 1. Regarding Claim 16: Claim 16 is analogous to claim 3. Regarding Claim 17: Claim 17 is analogous to claim 4. Regarding Claim 20: Claim 20 is analogous to claim 7. Claims 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hashem (“Optimal Linear Combinations of Neural Networks”), referred to as Hashem in this document, and further in combination with Samdani et al (“Domain Adaptation with Ensemble of Feature Groups”), referred to as Samdani in this document, and further in combination with Yeo et al (US 20220036152 A1), referred to as Yeo in this document, and further in combination with Meng et al (US 20200334538 A1), referred to as Meng in this document. Regarding Claim 11: The apparatus of claim 10 is taught by Hashem, Yeo, and Samdani Hashem teaches: wherein the multilayer neural network learning is based on a weighted sum of the one or more neural networks [Hashem 3. Linear Combinations of Neural Networks page 3]: “Consider a multi-input single-output mapping approximated by a trained NN. A trained NN accepts a vector-valued input x and returns a scalar output (response) y(x) [wherein the multilayer neural network learning is based on a weighted sum of the one or more neural networks]. The approximation error is δ(x) = r(x) − y(x), where r(x) is the response of the real system (true response) for x.” Where additional support for multilayer neural networks is given in [Hashem Introduction page 2]: “The class of neural networks investigated here is the class of multilayer feedforward networks. No further assumptions regarding the network architecture or the learning method are needed” Hashem does not explicitly teach: and a cross entropy loss function of GT- Label Meng teaches: and a cross entropy loss function of GT- Label [Meng 0027]: “where 0≤λ≤1 is the weight for the class posteriors and <custom character> is the indicator function which equals to 1 if the condition in the squared bracket is satisfied and 0 otherwise. Note that the interpolated T/S learning becomes soft T/S when λ=1.0 and becomes standard cross-entropy [and a cross entropy loss function of GT- Label] training with hard labels when λ=0.0. Although interpolated T/S compensates for the imperfection in knowledge transfer, the linear combination of soft and hard labels destroys the correct relationships among different classes embedded naturally in the soft class posteriors and deviates the student model parameters from the optimal direction. Moreover, the search for the best student model is subject to the heuristic tuning of λ between 0 and 1.” Further support for relation with ground truth is given in [Meng 0023]: “One shortcoming of T/S learning is that a teacher model, not always perfect, sporadically makes incorrect predictions that mislead the student model toward a suboptimal performance. In such a case, it may be beneficial to utilize hard labels of the training data to alleviate this effect. Some approaches use an interpolated T/S learning called knowledge distillation, in which a weighted sum of the soft posteriors and the one-hot hard label is used to train the student model. One issue is that the simple linear combination with one-hot vectors destroys the relationships among different classes embedded naturally in the soft posteriors produced by the teacher model. Moreover, proper setting of the interpolation weight with a fixed value is known to be critical and it varies with the adaptation scenarios and the qualities of the teacher and ground truth labels” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Meng. Hashem and Meng are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Meng in order to be able to find a better model as the method acts to aid training ([Meng 0027]: “where 0≤λ≤1 is the weight for the class posteriors and <custom character> is the indicator function which equals to 1 if the condition in the squared bracket is satisfied and 0 otherwise. Note that the interpolated T/S learning becomes soft T/S when λ=1.0 and becomes standard cross-entropy training with hard labels when λ=0.0… Moreover, the search for the best student model is subject to the heuristic tuning of λ between 0 and 1.”) Regarding Claim 12: The apparatus of claim 10 is taught by Hashem, Yeo, and Samdani Hashem does not explicitly teach: wherein the multilayer neural network learning is performed based on a knowledge distillation method Meng teaches: wherein the multilayer neural network learning is performed based on a knowledge distillation method [Meng 0023]: “One shortcoming of T/S learning is that a teacher model, not always perfect, sporadically makes incorrect predictions that mislead the student model toward a suboptimal performance. In such a case, it may be beneficial to utilize hard labels of the training data to alleviate this effect. Some approaches use an interpolated T/S learning called knowledge distillation [wherein the multilayer neural network learning is performed based on a knowledge distillation method], in which a weighted sum of the soft posteriors and the one-hot hard label is used to train the student model. One issue is that the simple linear combination with one-hot vectors destroys the relationships among different classes embedded naturally in the soft posteriors produced by the teacher model. Moreover, proper setting of the interpolation weight with a fixed value is known to be critical and it varies with the adaptation scenarios and the qualities of the teacher and ground truth labels.” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Meng. Hashem and Meng are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Meng in order to be able to create a better model, such as a student model that is smarter than the teacher model ([Meng 0024]: “Some embodiments described herein utilize a conditional T/S learning scheme where a student model becomes smart so that it can criticize the knowledge imparted by the teacher model to make better use of the teacher and the ground truth. At the initial stage, when the student model is very weak, it may blindly follow all knowledge infused by the teacher model and use the soft posteriors as the sole training targets. As the student model grows stronger, it may begin to selectively choose the learning source from either the teacher model or the ground truth labels conditioned on whether the teacher's prediction coincides with the ground truth. That is, the student model may learn exclusively from the teacher when the teacher makes correct predictions on training samples, and otherwise from the ground truth when the teacher is wrong. With conditional T/S learning, the student makes good use of rich and correct knowledge encompassed by the teacher yet avoids receiving inaccurate knowledge generated by the teacher. Another advantage of the conditional T/S learning over the conventional T/S learning is that it forgoes tuning the interpolation weight between two knowledge sources.” ) Claims 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hashem (“Optimal Linear Combinations of Neural Networks”), referred to as Hashem in this document, and further in combination with Samdani et al (“Domain Adaptation with Ensemble of Feature Groups”), referred to as Samdani in this document, and further in combination with Yeo et al (US 20220036152 A1), referred to as Yeo in this document, and further in combination with Liang et al (“Understanding Mixup Training Methods”), referred to as Liang in this document. Regarding Claim 13: The apparatus of claim 10 is taught by Hashem, Yeo, and Samdani Hashem does not explicitly teach: wherein the learning data is generated by a mixup method of adjusting a ratio of data to the prototype domain Liang teaches: wherein the learning data is generated by a mixup method of adjusting a ratio of data to the prototype domain [Liang 3. Method A. General Mixup page 3]: “In order to explore the impact of the mixing of images and their labels on the performance of the network [wherein the learning data is generated by a mixup method of adjusting a ratio of data to the prototype domain as Liang is teaching the mixup method and related information], we either interpolate the images or interpolate their labels, or interpolate them at the same time. The Uniform distribution can better control the range of λ compared to the Beta distribution, and has the same effect on some datasets. Therefore, the Uniform distribution is adopted to control the sampling of λ . We use λx to represent the mixing ratio of two samples x for λ∈ Uniform(λ1,λ2) , and Rl to represent the mixing ratio of two labels y for λ∈ Uniform(λ1,λ2) , where 0≤λ1≤λ≤λ2≤1 . When λx=λl , we denote them as λ .” One of ordinary skill in the art, prior to the effective filing date, would have been motivated to combine Hashem and Meng. Hashem and Meng are in the same field of endeavor of machine learning. One of ordinary skill in the art would have been motivated to combine Hashem and Meng in order to be able to prevent overfitting or improve generalization ([Liang Introduction page 1]: “In SamplePairing [13] and mixup [14], a method of training a neural network using two samples simultaneously is proposed. SamplePairing randomly picks one sample in the training set to add to the original sample and uses the original sample’s label to train the network. The mixup uses a random value to weight the two samples and their corresponding labels. All of the above methods have some effect of data augmentation and regularization, and they can achieve better generalization performance than Empirical Risk Minimization (ERM) [15].”) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chu et al (US 20210374617 A1) is considered relevant art, as Chu et al discusses the use of weighted averaging but utilizing coefficients, aka weighted aggregation, in paragraphs 74-75. Thus Chu et al is an example of some of the ideas in the current application. Gupta et al (“Stochastic Weight Ave raging in Parallel: Large-batch Training that Generalizes Well”) is relevant art that discusses a method of combining model parameters in a method called SWAP or Stochastic Weight Averaging in Parallel. Thus Gupta et al is relevant to the current application in that Gupta et al discusses the creation of a new “final” model utilizing the combination of existing models weights. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRISTOPHER D DEVORE whose telephone number is (703)756-1234. The examiner can normally be reached Monday-Friday 7:30 am - 5 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.D.D./Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Dec 28, 2022
Application Filed
Jan 30, 2026
Non-Final Rejection mailed — §103
Apr 30, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705478
IMPLICIT CURRICULUM LEARNING
5y 1m to grant Granted Aug 11, 2026
Patent 12530603
OBTAINING AND UTILIZING FEEDBACK FOR AGENT-ASSIST SYSTEMS
4y 7m to grant Granted Jan 20, 2026
Patent 12505355
GENERAL FORM OF THE TREE ALTERNATING OPTIMIZATION (TAO) FOR LEARNING DECISION TREES
4y 0m to grant Granted Dec 23, 2025
Patent 12468978
Reinforcement Learning In A Processing Element Method And System Thereof
4y 0m to grant Granted Nov 11, 2025
Patent 12412069
COOKIE SPACE DOMAIN ADAPTATION FOR DEVICE ATTRIBUTE PREDICTION
3y 10m to grant Granted Sep 09, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
57%
Grant Probability
93%
With Interview (+36.1%)
4y 2m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 21 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month