Prosecution Insights
Last updated: August 15, 2026
Application No. 17/064,706

Method and System For Sharing Meta-Learning Method(s) Among Multiple Private Data Sets

Non-Final OA §103
Filed
Oct 07, 2020
Examiner
SMITH, KEVIN LEE
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Cognizant Technology Solutions U S Corporation
OA Round
5 (Non-Final)
38%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
57%
With Interview

Examiner Intelligence

Grants only 38% of cases
38%
Career Allowance Rate
52 granted / 138 resolved
-17.3% vs TC avg
Strong +19% interview lift
Without
With
+19.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
31 currently pending
Career history
184
Total Applications
across all art units

Statute-Specific Performance

§101
31.2%
-8.8% vs TC avg
§103
39.8%
-0.2% vs TC avg
§102
10.9%
-29.1% vs TC avg
§112
13.5%
-26.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 138 resolved cases

Office Action

§103
DETAILED ACTION 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination 2. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant’s submission filed 12 May 2026 [hereinafter Response] has been entered, where: Claims 1, 6, 12, 17, and 18 have been amended. Claims 3, 5, 10, 14, and 16 have been cancelled. Claims 1, 2, 4, 6-9, 11-13, 15, 17 and 18 are pending. Claims 8, 9, and 11 are rejected. Claims 1, 2, 4, 6, 7, 12, 13, 15, 17, and 18 are allowed. Claim Rejections - 35 U.S.C. § 103 3. The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 4. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. § 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 5. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention. 6. Claims 8, 9, and 11 are rejected under 35 U.S.C. § 103 as being unpatentable over US Published Application 20210004676 to Jaderberg et al. [hereinafter Jaderberg] in view of US Published Application 20210182660 to Amirguliyev et al. [hereinafter Amirguliyev] and Oehmcke et al., “Knowledge Sharing for Population Based Neural Network Training,” Springer (2018) [hereinafter Oehmcke]. Regarding claim 8, Jaderberg teaches [a] process for evolving and providing access to a teacher neural network model trained on private data for use in evolving a child neural network model for solving a predetermined problem or making a prediction (Jaderberg ¶ 0039 teaches [t]he method (that is, process) enables neural network training to be carried out more efficiently and to produce a better trained neural network), the process comprising: creating and training a teacher neural network model by a first system including a first server (Jaderberg ¶ 0005 teaches [d]uring the training of the neural network, the system maintains a plurality of candidate neural networks (that is, creating and training a teach model); Jaderberg ¶ 0039 teaches that the training is such that it may be performed by a distributed system and the described approach only requires values of parameters, hyperparameters, and quality measures to be communicated between candidates (that is, the “distributed system” provides the first system including a first server), . . . ; determining by the first system a set of performance metrics for each of the teacher neural network model (Jaderberg ¶ 0051 teaches [d]uring training, the network parameters, the hyperparameters, and the quality measure for a candidate neural network are updated in accordance with training operations, including an iterative training process (that is, “during training, . . . the quality measure . . . [is] updated” is determining by the first system a set of performance metrics for each of the teacher model)), . . . ; providing access to the trained teacher neural network model in a commonly accessible area (Jaderberg ¶ 0049 teaches system 100 can output (e.g., by outputting to a user device or by storing in memory accessible to the system) the trained values of the network parameters of the trained neural network 150 for later use in processing inputs using the trained neural network 150 (that is, providing access to the trained teacher model in a commonly accessible area)), . . . ; evolving a . . . neural network model in accordance with one or more domain factors (Jaderberg ¶ 0055 teaches system 100 trains each candidate neural network 120A-N by repeatedly performing iterations of an iterative training process to determine updated network parameters for the respective candidate neural network (that is, the “repeatedly performing iterations” is evolving a . . . model); Jaderberg ¶ 0119 teaches a supervised learning task is specified as the machine learning task for the system (that is, the “supervised learning task” is one or more domain factors); e.g., Jaderberg teaches that these one or more domain factors include a deep reinforcement machine learning task (Jaderberg ¶ 0111), a machine translation learning task (Jaderberg ¶ 0119), and a set of training operations for a populations of pairs of candidate and discriminator neural networks of a GAN task (Jaderberg ¶ 0126)), the evolving including: (i) creating by a second subsystem, including a second server, a first population of candidate . . . neural network models and assigning a unique candidate identifier to each of the candidate . . . neural network models in the first population (Jaderberg, Fig. 1, teaches an initial population P that assigns a unique candidate identifier to each of the candidate student models in the first population [Examiner annotations in dashed-line text boxes]: PNG media_image1.png 624 559 media_image1.png Greyscale Jaderberg ¶ 0045 teaches example population based neural network training system (that is, creating by a second subsystem, including a second server, a first population of candidate . . . models and assigning a unique candidate identifier to each of the candidate . . . models in the first population); Jaderberg ¶ 0149 teaches [i]mplementations of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server) (that is, a second subsystem, including a second server); see also Jaderberg ¶ 0142, which teaches an “engine” is a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers); (ii) transmitting the first population of candidate student neural network models with assigned candidate identifiers, to a third subsystem, including a third server (Jaderberg ¶ 0039 teaches that the training is such that it may be performed by a distributed system and the described approach only requires values of parameters, hyperparameters, and quality measures to be communicated between candidates (that is, the “distributed system” provides the third subsystem, including a third server); Jaderberg ¶ 0072 teaches [a] population based neural network training system 100 benefits from local optimization by executing, asynchronously and in parallel, an iterative training process for each candidate neural network 120A-N in the population (that is, “local optimization” accessing the population by transmitting the first population of candidate . . . models)), . . . ; (iii) training by the third subsystem the first population of candidate . . . neural network models (Jaderberg ¶ 0005 teaches [d]uring the training of the neural network, the system maintains a plurality of candidate neural networks (that is, training . . . the first population of candidate . . . models); Jaderberg ¶ 0039 teaches that the training is such that it may be performed by a distributed system and the described approach only requires values of parameters, hyperparameters, and quality measures to be communicated between candidates (that is, the “distributed system” provides the third subsystem)) . . . ; (iv) determining by the third subsystem a set of performance metrics for each of the candidate student neural network models (Jaderberg ¶ 0051 teaches [d]uring training, the network parameters, the hyperparameters, and the quality measure for a candidate neural network are updated in accordance with training operations, including an iterative training process (that is, “during training, . . . the quality measure . . . [is] updated” is determining by the second subsystem a set of performance metrics for each of the candidate teacher models)), wherein the set of performance metrics is indicative of a fitness of each of the candidate teacher models for solving the predetermined problem or making a prediction (Jaderberg ¶ 0111 teaches “a reinforcement learning task is specified as the machine learning task for the system. Each candidate neural network receives inputs and generates outputs that conform to a deep reinforcement learning task [(that is, “task” is for solving the predetermined problem or making a prediction)]”; Jaderberg ¶ 0113 teaches “eval(•): [t]he system updates the quality measure of a candidate neural network based on the mean value of a pre-determined number of previous episodic rewards (e.g., 10 episodic rewards). Specifically, the quality measure of the candidate neural network is the mean value of the predetermined number of previous episodic rewards. The candidate neural network with the highest mean episodic reward [(that is, each of the candidate teacher models)] has the highest quality measure and is considered the ‘best’ in terms of measured fitness [(that is, the set of performance metrics is indicative of a fitness of each of the candidate teacher models for solving the predetermined problem or making a prediction)]”) . . . ; (v) providing the set of performance metrics for each of the candidate . . . neural network models in accordance with assigned candidate identifier to the second subsystem (Jaderberg ¶ 0054 teaches that [t]he quality measure for a candidate neural network is a measure of the performance of the candidate neural network on the particular machine learning task (that is, providing the set of performance metrics for each of the candidate . . . models in accordance with assigned candidate identifier to the second subsystem)); (vi) creating a next population of candidate student neural network models from the sets of performance metrics for each of the candidate . . . neural network models from the third subsystem (Jaderberg ¶ 0009 teaches system determines a new quality measure for the candidate neural network in accordance with the new values of the network parameters for the candidate neural network and the new values of the hyperparameters for the candidate neural network and updates the maintained data for the candidate neural network to specify the new values of the hyperparameters, the new values of the network parameters, and the new quality measure; Jaderberg ¶ 0010 teaches [a]fter repeatedly performing the set of training operations, the system selects the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures for the candidate neural networks after the training operations have repeatedly been performed (that is, a population of “updated candidate neural networks” is creating a next population of candidate teach models from the sets of performance metrics for each of the candidate . . . models from the third subsystem)); (vii) repeating steps (ii) to (vi) for multiple generations until an optimized candidate . . . neural network model is determined in accordance with one or more predetermined conditions (Jaderberg ¶ 0136 teaches [w]hen training is over, the system generates data specifying a trained neural network by selecting the candidate neural network from the population with the highest quality measure (shown in line 16 in TABLE 1) (that is, a best candidate . . . model is determined in accordance with a predetermined condition); Jaderberg ¶ 0136 teaches the system checks if the termination criteria are satisfied (shown as a conditional statement in line 6 in TABLE 1 executing ready(p, t, P)) (that is, the “termination criteria” is a predetermined condition)); providing access to the optimized candidate . . . neural network model in the commonly accessible area (Jaderberg ¶ 0049 teaches system 100 can output (e.g., by outputting to a user device or by storing in memory accessible to the system) the trained values of the network parameters of the trained neural network 150 for later use in processing inputs using the trained neural network 150 (that is, providing access to the best candidate. . . model in the commonly accessible area)); Jaderberg ¶ 0066 teaches [an] optimal candidate neural network selected is sometimes referred to in this specification as the “best” candidate neural network 120A-N in the population. A candidate neural network that has a higher quality measure than another candidate neural network is considered “better” than the other candidate neural network (that is, the best candidate . . . model)), . . . ; and * * * Though Jaderberg teaches population based neural network training system that benefits from local optimization, Jaderberg, however, does not explicitly teach – * * * [creating and training] . . . , wherein the teacher neural network model is trained on a first secure data set; [determining] . . . , wherein the set of performance metrics does not include secure data from the first secure data set; [providing access] . . . , wherein the first subsystem and the commonly accessible area are separated by a first firewall which blocks access to the first secure data set; * * * [(ii) transmitting] . . . , wherein the second server and the third server are separated by a second firewall which blocks access to a second secure data set; [(iii) training] . . . against the second secure data set located behind the second firewall; [(iv) determining] . . . , wherein the set of performance metrics does not include secure data from the second secure data set; * * * [providing access] . . . , wherein the access does not include access to secure data from the first or second secure data sets, and further wherein the second subsystem and the commonly accessible area are separated by a third firewall; and * * * But Amirguliyev teaches - * * * [creating and training] . . . , wherein the teacher neural network model is trained on a first secure data set (Amirguliyev ¶ 0067 teaches “the master device 210 is configured to instantiate the first version of the neural network model 262 [(that is, the teacher neural network)] using the parameter data from the storage device 260. This includes a process similar to that described for the second version of the neural network model 170 in the example of FIG. 1. The first version of the neural network model 262 may be instantiated using the first configuration data (CD1) 280 or similar data. The master device 210 is then configured to train the first version of the neural network model 262 using data from the master data source 264 [(that is, wherein the teacher neural network model is trained on a first secure data set)]”); [determining] . . . , wherein the set of performance metrics does not include secure data from the first secure data set (Amirguliyev ¶ 0067 teaches “The first version of the neural network model 262 may be instantiated using the first configuration data (CD1) 280 or similar data [(that is, the “configuration data (CD1)” is wherein the set of performance metrics does not include secure data from the first secure data set)]”); [providing access] . . . , wherein the first subsystem and the commonly accessible area are separated by a first firewall which blocks access to the first secure data set (Amirguliyev, Fig. 2, teaches a firewall separating servers [Examiner annotations in dashed-line text boxes]:” PNG media_image2.png 718 664 media_image2.png Greyscale Amirguliyev ¶ 0066 teaches “the local area network 212 has a firewall that prevents or controls access to at least devices on the master local area network 212. The master device 210 also has a storage device 260 to store parameter data for an instantiated first version of a neural network model 262 [(that is, wherein the first subsystem and the commonly accessible area are separated by a first firewall which blocks access to the first secure data set)]”); * * * [(ii) transmitting] . . . , wherein the second server and the third server are separated by a second firewall which blocks access to a second secure data set (Amirguliyev, Fig. 2, teaches a firewall separating servers [Examiner annotations in dashed-line text boxes]:” PNG media_image2.png 718 664 media_image2.png Greyscale Amirguliyev ¶ 0066 teaches “the local area network 212 has a firewall that prevents or controls access to at least devices on the master local area network 212. The master device 210 also has a storage device 260 to store parameter data for an instantiated first version of a neural network model 262 [(that is, the first server and the second server are separated by a first firewall which blocks access to a first secure data set)]”); [(iii) training] . . . against the second secure data set located behind the first firewall (Amirguliyev ¶ 0068 “Following training on data from the slave data source 230, the slave device 220 stores an updated set of parameters in the storage device 272 and use these to generate a second configuration data (CD2) 290 that is sent to the master device 210. The master device 210 receives the second configuration data 290 and uses it to update the parameter data stored in the storage device 260 [(that is, [(iii) training] . . . against a first secure data set located behind the first firewall)]”); [(iv) determining] . . . , wherein . . . the set of performance metrics does not include secure data from the second secure data set (Amirguliyev ¶ 0100 teaches “[o]ne or more performance metrics may be determined and compared to predefined thresholds; performance below a threshold may lead to exclusion or down-weighting of a particular instance [(that is, via the firewall, it is inherent that the set of performance metrics does not include secure data from the second secure data set)]”); * * * [providing access] . . . , wherein the access does not include access to secure data from the first or second secure data sets, and further wherein the second subsystem and the commonly accessible area are separated by a third firewall (Amirguliyev, Fig. 7, teaches a distributed training system 700 [Examiner annotations in dashed-line text boxes] PNG media_image3.png 714 520 media_image3.png Greyscale Amirguliyev ¶ 0088 teaches “a plurality of slave devices 722, 724, and 726 via one or more communication networks 750 . . . communicatively coupled to a respective slave data source 732, 734, and 736, which as per FIGS. 1 and 2 may be inaccessible by the master device 710, e.g. may be located on respective private networks 742, 744, and 746. The slave data sources 732, 734 and 736 may also be inaccessible by other slave devices, e.g. the slave device 724 may not be able to access the slave data source 732 or the slave data source 736. Each slave device 722, 724, and 726 receives the first configuration data 780 from the master device 710 as per the previous examples and generates respective second configuration data (CD2) 792, 794, and 796, which is communicated back to the master device 710 via the one or more communication networks 750. The set of second configuration data 792, 794, and 796 from the plurality of slave devices 722, 724, and 726, respectively, may be used to update parameters for the first version of the neural network model [(that is, wherein the access does not include access to secure data from the first or second secure data sets, and further wherein the second subsystem and the commonly accessible area are separated by a third firewall)]”); * * * Though Jaderberg and Amirguliyev teach the features of population-based training in a distributed environment with a firewall, the combination of Jaderberg and Amirguliyev, however, does not explicitly teach – * * * and training the optimized candidate student neural network model in accordance with operation by the trained teacher neural network model on one or more additional datasets to solve the predetermined problem or make the prediction, wherein the one or more optimized candidate student neural network model is smaller than the optimized candidate teacher neural network model. But Oehmcke teaches that “knowledge distilling” is the distilling of knowledge for neural networks in which a complex model is trained, and then let it be the teacher for a simpler model, the student (that is, candidate student model). The student model is then able to achieve nearly the same performance as the complex one, which was not possible when training the simple [student] model without the teacher. (Oehmcke at p. 265, “4.1 Distilling Knowledge,” first paragraph). Oehmcke also teaches - and training the optimized candidate student neural network model in accordance with operation by the trained teacher neural network model on one or more additional datasets (Oehmcke at p. 259, “1. Introduction,” second paragraph, teaches a novel extension to [population based training (PBT)] by enabling knowledge sharing across generations. We adapt a knowledge distilling strategy, which is inspired by Hinton et al. [10], where the knowledge about the training data of the best individuals is stored separately and fed back to all individuals via the loss function) to solve the predetermined problem or make the prediction (Oehmcke teaches extensions to population-based training with knowledge distilling in Algorithm 1 (Examiner annotations in dashed-line text boxes): PNG media_image4.png 543 799 media_image4.png Greyscale Oehmcke at p. 260, “Knowledge Sharing,” first paragraph, teaches extensions to the [population-based training] with knowledge distilling [as] highlighted with [shaded boxes] in Algorithm 1. The teacher output T = {t1, . . . , tn} ∈ ℝc is initialized with the one-hot-encoded class targets of the true training targets Ytrain (Line 1)). During the evolutionary process, the best models are allowed to contribute to T through the teach-function (Line 13). We implement this teach-function by replacing 20% of the teacher output T with the predicted probability if the individual is from the upper 20% of the population P regarding the fitness p (that is, the one or more additional datasets by the best candidate teacher model). Depending on the population size, we are able to replace the original targets from Y in a few generations and introduce updates from generations continuously (that is, “generations” is evolving); Oehmcke at p. 261, “2.2 Knowledge Sharing,” second & third paragraph, teaches [the] combination of cross entropy and Kullback-Leibler divergence ensures that the models can learn the true labels, while also utilizing the already acquired knowledge of the population. The trade-off parameter α is added to the hyperparameter settings h. To compare the output distributions of the teacher and the individuals (that is, the candidate student models), we employ the Kullback–Leibler divergence DKL inspired by other distilling approaches [23]. The one-hot encoding of the true target as initialization results in a loss function equal to only using cross entropy since the Kullback-Leibler divergence is approximately equal to the cross entropy when all-except-one class probabilities are zero. There are similarities to generative adversarial networks (GANs) [7], where the generator is similar to the teacher (that is, the teacher model) and discriminator is similar to the student (that is, the student model)), wherein the one or more optimized candidate student neural network model is smaller than the optimized candidate teacher neural network model (Oehmcke at p. 265, “4.1 Distilling Knowledge,” first paragraph, teaches “They trained a complex model and let it be the teacher for a simpler model, the student. The student model is then able to achieve nearly the same performance as the complex one, which was not possible when training the simple model without the teacher). Jaderberg, Amirguliyev, and Oehmcke are from the same or similar field of endeavor. Jaderberg teaches population based training of neural networks. Amirguliyev teaches distributed training of a neural network model where a master device has a first version of the neural network model and a slave device is communicatively coupled to a first data source and the master device, and the first data source is inaccessible by the master device. Oehmcke teaches extensions to population-based training to provide knowledge distilling (or sharing). Thus, it would have been obvious to a person having ordinary skill in the art as of the effective filing date of the Applicant’s invention to modify the combination of Jaderberg and Amirguliyev pertaining to distributed neural network training having a firewall with the population-based training extended to knowledge distillation of Oehmcke. The motivation to do so is for a novel knowledge distilling scheme where only the best individuals of the population are allowed to share part of their knowledge about the training data with the whole population to embrace the idea of randomness between the models, rather than avoiding it, because the resulting diversity of models is important for the population’s evolution. (Oehmcke, Abstract). Regarding claim 9, the combination of Jaderberg, Amirguliyev, and Oehmcke teaches all of the limitations of claim 8, as described above in detail. Jaderberg teaches - wherein the one or more domain factors are selected from the group consisting of: domain constraints, known domain parameters and formatting rules for a specific representation of each of the candidate teacher neural network models and candidate student neural network models (Jaderberg ¶ 0012 teaches [t]he hyperparameters can include values that impact how the values of the network parameters are updated by the training process e.g., the learning rate or other update rule (that is, formatting rules) that defines how the gradients determine at the current training iteration are used to update the network parameter values, objective function values, e.g., entropy cost, weights assigned to various terms of the objective function, and so on; Jaderberg ¶ 0126 teaches, under a set of training operations for a population of pairs of candidate generator and discriminator neural networks of a GAN task (that is, one or more domain factors); Jaderberg ¶ 0130 teaches [i]f the pair of candidate neural networks is ranked below a certain threshold (e.g., bottom 20% of all pairs of candidate neural networks), the system samples another pair of candidate neural networks ranked above a certain threshold (e.g., top 20% of all pairs candidate neural networks) and sets the new network parameters and new hyperparameters of the generator candidate neural network to the values of the network parameters and hyperparameters of the sampled generator candidate neural network (that is, in relating to the “GAN task,” these are known domain parameters) and sets the new network parameters and new hyperparameters of the discriminator candidate neural network to the values of the network parameters and hyperparameters of the sampled discriminator candidate neural network (that is, for a specific representation of each of the candidate individuals)). Regarding claim 11, the combination of Jaderberg, Amirguliyev, and Oehmcke teaches all of the limitations of claim8, as described above in detail. Amirguliyev teaches - wherein the first subsystem and the commonly accessible location are separated by a second firewall (Amirguliyev, Fig. 7, teaches a distributed training system 700 [Examiner annotations in dashed-line text boxes] PNG media_image3.png 714 520 media_image3.png Greyscale Amirguliyev ¶ 0088 teaches “a plurality of slave devices 722, 724, and 726 via one or more communication networks 750 . . . communicatively coupled to a respective slave data source 732, 734, and 736, which as per FIGS. 1 and 2 may be inaccessible by the master device 710, e.g. may be located on respective private networks 742, 744, and 746. The slave data sources 732, 734 and 736 may also be inaccessible by other slave devices, e.g. the slave device 724 may not be able to access the slave data source 732 or the slave data source 736. Each slave device 722, 724, and 726 receives the first configuration data 780 from the master device 710 as per the previous examples and generates respective second configuration data (CD2) 792, 794, and 796, which is communicated back to the master device 710 via the one or more communication networks 750. The set of second configuration data 792, 794, and 796 from the plurality of slave devices 722, 724, and 726, respectively, may be used to update parameters for the first version of the neural network model [(that is, wherein the access does not include access to secure data from the first or second secure data sets, and further wherein the second subsystem and the commonly accessible area are separated by a third firewall)]”) Allowable Subject Matter Claims 1, 2, 4, 6, 7, 12, 13, 15, 17, and 18 are allowed. Response to Arguments 7. Examiner has fully considered Applicant’s arguments, and responds below accordingly. Claim Rejections – 35 U.S.C. §103 8. Applicant argues “In the present Office Action and the Advisory Action, the Examiner maintains that the claims are unpatentable over the combination of Jaderberg, Amirguliyev and Oehmcke. In the Advisory Action, the Examiner maintained their position that the differentiating elements were only supported by the preamble - not the body of the claim - and thus were not being afforded patentable weight. The Examiner suggests that amending the independent claims to include the elements, such as those in Claim 5, directed to the process for evolving a child neural network model for solving a predetermined problem or making a prediction would advance prosecution. Accordingly, claims 1 and 12 have been amended herein. Independent claim 8 already includes these elements. The Applicant submits that with the positive recitation of these elements within the body of the claims, the combination of elements in the claims is patentable over the combination of Jaderberg, Amirguliyev and Oehmcke since the combination is missing critical elements of the claims as has been particularly pointed out in the prior responses as follows.” (Response at p. 10). Examiner’s Response: With regard to claims 1, 2, 4, 6, 7, 12, 13, 15, 17, and 18, Examiner finds Applicant’s amendments and arguments thereto persuasive, and accordingly, withdraws the rejection under Section 103 to these claims. With respect to claims 8, 9, and 11, Examiner addressed the instant arguments in the Advisory Action dated 04 May 2026, which is incorporated herein. Conclusion 9. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: (US Published Application 20200034702 to Fukuda et al.) teaches that a student neural network may be trained by a computer-implemented method, including: selecting a teacher neural network among a plurality of teacher neural networks, inputting an input data to the selected teacher neural network to obtain a soft label output generated by the selected teacher neural network, and training a student neural network with at least the input data and the soft label output from the selected teacher neural network. (US Published Application 20030233335 to Mims) teaches A student neural network that is capable of receiving a series of tutoring inputs from one or more teacher networks to generate a student network output that is similar to the output of the one or more teacher networks. The tutoring inputs are repeatedly processed by the student until, using a suitable method such as back propagation of errors, the outputs of the student approximate the outputs of the teachers within a predefined range. Once the desired outputs are obtained, the weights of the student network are set. Using this weight set the student is now capable of solving all of the problems of the teacher networks without the need for adjustment of its internal weights. If the user desires to use the student to solve a different series of problems, the user only needs to retrain the student by supplying a different series of tutoring inputs. (Hongpeng Zhu, Online Meta-Learning Firewall to Prevent Phishing Attacks,” Neural Computing and Applications (June 2020)) teaches It is a highly innovative and fully automated active safety tool that uses a long short-term memory meta-learner algorithm. This method can learn to efficiently classify using a small number of samples. At the same time, it can converge with a fairly small number of steps. The proposed system is an improvement on the k-nearest neighbor with self-adjusting memory algorithm, which is inspired by the model of short and long-term memory. The purpose of the system is to understand the nature of an unknown situation and to classify it, based on the most relevant characteristics that come directly from the unknown environment. 10. Any inquiry concerning this communication or earlier communications from the Examiner should be directed to KEVIN L. SMITH whose telephone number is (571) 272-5964. Normally, the Examiner is available on Monday-Thursday 0730-1730. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, KAKALI CHAKI can be reached on 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.L.S./ Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Show 6 earlier events
May 04, 2025
Response after Non-Final Action
Jun 10, 2025
Non-Final Rejection mailed — §103
Oct 08, 2025
Response Filed
Jan 27, 2026
Final Rejection mailed — §103
Apr 21, 2026
Response after Non-Final Action
May 12, 2026
Request for Continued Examination
May 16, 2026
Response after Non-Final Action
May 26, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664451
SYSTEM AND METHOD FOR GENERATING A PREDICTIVE MODEL
6y 3m to grant Granted Jun 23, 2026
Patent 12657425
DYNAMIC CACHE MANAGEMENT IN BEAM SEARCH
5y 3m to grant Granted Jun 16, 2026
Patent 12591815
METHOD AND SYSTEM FOR UPDATING MACHINE LEARNING BASED CLASSIFIERS FOR RECONFIGURABLE SENSORS
4y 10m to grant Granted Mar 31, 2026
Patent 12585917
REINFORCEMENT LEARNING USING ADVANTAGE ESTIMATES
4y 0m to grant Granted Mar 24, 2026
Patent 12547759
PRIVACY PRESERVING MACHINE LEARNING MODEL TRAINING
5y 6m to grant Granted Feb 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
38%
Grant Probability
57%
With Interview (+19.4%)
4y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 138 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month