Prosecution Insights
Last updated: October 02, 2026
Application No. 17/471,118

RANDOMIZED PARAMETER SETTING FOR MODEL TRAINING

Final Rejection §103§112
Filed
Sep 09, 2021
Priority
Sep 11, 2020 — provisional 63/077,239
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Actapio Inc.
OA Round
4 (Final)
51%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
41 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on June 4, 2026, in which claims 1, 8, and 15-18 are currently amended. Claims 1-18 are currently pending. Response to Arguments The previous rejections to claims 1-18 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claim 18 under 35 U.S.C. 101 based on amendment have been considered and are persuasive. The previous rejections to claims 1-18 under 35 U.S.C. § 101 are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-18 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 1, 16, 17, and 18, "to reduce processing time for the prediction process" is indefinite. It would be unclear to one of ordinary skill in the art what the processing time for the prediction process was reduced relative to: other models, a predetermined threshold, earlier training epochs of the same model, simply having low latency, or something else altogether. As these interpretations are contradictory the scope of the claim cannot reasonably be determined. In the interest of further examination having low latency is interpreted as synonymous with having reduced processing time. Regarding claim 9, "wherein the predetermined evaluation conditions comprise a change in the evaluation value satisfies a predetermined mode" is grammatically indefinite. After "comprise," a noun phrase is expected such as "the predetermined evaluation conditions comprise a change in the evaluation", however, the sentence continues with another finite verb "satisfies a predetermined mode" such that "satisfies" has no clear grammatical role. It could mean the change in evaluation value satisfies the predetermined mode, the evaluation conditions include a requirement that the change satisfies the predetermined mode, or something else altogether. Since these interpretations are contradictory the scope of the claim cannot reasonably be determined. In the interest of further examination the claim is interpreted as "the predetermined evaluation conditions comprise a condition that a change in the evaluation value satisfies a predetermined mode". The remaining claims are rejected with respect to their dependence on the rejected claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 16, and 17 are rejected under U.S.C. §103 as being unpatentable over the combination of Fu (“Enhancing Model Parallelism in Neural Architecture Search for Multidevice System”, 2020) and Mahyar (US10930263B1). PNG media_image1.png 630 756 media_image1.png Greyscale FIG. 1 of Fu Regarding claim 1, Fu teaches A learning apparatus comprising: a processor, ([p. 47] "NAS techniques 4–6 target to identify efficient models on single CPU/GPU systems" [p. 48] "multiple devices are available for parallelism" [p. 51] "We run ColocNAS on different device configurations […] Figure 3 shows the effectiveness of our kNN based latency predictor. The latency is profiled in search space ColocNAS-SP-4 with ImageNet input (batch size is 32) on 4 Tesla K80 GPUs. We use 5% of the profiled data for testing. The profiling takes 2.5 days on 8 GPUs") the processor is configured to: generate a plurality of models each having different parameters, ([p. 48] "A is a new search space proposed in ColocNAS, a 2 A is a set of continuous variables that specify a possible architecture, wa is the weight parameter of the network. […] Mi <- Random_Generate(Nt, Ops, Nt, Nc)" [p. 50] "We randomly generate architectures and profile their latency values on the target hardware" [p. 50] "the architecture parameter a and the weights wa") wherein the parameters comprise weights ([p. 48] "A is a new search space proposed in ColocNAS, a 2 A is a set of continuous variables that specify a possible architecture, wa is the weight parameter of the network" [p. 50] "the architecture parameter a and the weights wa") to learn features of a part of predetermined learning data;([p. 48 Algorithm 1] "Mout <- Training_Model(March)" [p. 50] "after the model is searched, we train the model from scratch" [p. 7] "After the architecture is search, we train the weight parameters of the resulting network for 1500 epochs using a batch size of 96 and Adam optimizer" batch interpreted as a part of predetermined learning data) select one of the plurality of models based on model accuracy ([p. 48] "A promising architecture yields small latency and high task accuracy" [p. 50] "we sample N neural network architectures arch1;arch2;:::;archN. The expectation value of the latency variable is then approximated as [...] ColocNAS aims to find a network architecture with low latency and high task accuracy. As such, we define the loss function of ColocNAS’s online searching phase as follows: [See Eqn. 7 and 8]") train the selected model to learn features of the predetermined learning data. ([p. 48 Algorithm 1] "Mout <- Training_Model(March)" [p. 50] "after the model is searched, we train the model from scratch" [p. 7] "After the architecture is search, we train the weight parameters of the resulting network for 1500 epochs using a batch size of 96 and Adam optimizer") wherein a prediction process using the trained selected model includes a plurality of processes, ([p. 48] "Figure 2. Example of basic building blocks in three types of search space. (a) Layer-wise searched model. (b) Cell-based searched model. (c) New search space proposed in ColocNAS. The yellow circle refers to the concatenation node. bi n denotes the nth block in ith cell (i.e., layer)" blocks interpreted as processes which the prediction process using the trained selected model explicitly uses) and the learning apparatus decides, based on features of the trained selected model, ([p. 50] "We apply the RL-based method proposed by Mirhoseini et al.9 where the policy network is an attentional autoencoder that takes the NN graph as input") which arithmetic unit among a plurality of arithmetic units having different architectures is to execute each of the plurality of processes included in the prediction process, ([p. 47] "NAS techniques 4–6 target to identify efficient models on single CPU/GPU systems" [p. 50] "our expert placement policy uniformly distributes each concatenation node and the associated operations onto all existing devices [...] considering the heterogeneous property of the underlying computing platforms. To further optimize the latency of the searched architecture G, the last step of the ColocNAS leverages policy gradient for device placement on the given hardware") to reduce processing time for the prediction process.([p. 50] "ColocNAS aims to find a network architecture with low latency and high task accuracy."). However, Fu does not explicitly teach wherein the parameters comprise weights and biases; train each of the plurality of models. Mahyar, in the same field of endeavor, teaches wherein the parameters comprise weights and biases;([Col. 5 l. 15-20] "The hyper parameters of each neural network and the values of parameters following training (e.g., weights and biases of artificial neurons, forget gates, etc.)") train each of the plurality of models to learn features of a part of predetermined learning data ([Col. 5 l. 45-50] "Each voice synthesis model is trained" [Col. 12 l. 2-3] "Model selection at step 530 can optimize and prioritize such differences to prefer certain models over other models"). Fu as well as Mahyar are directed towards neural architecture search. Therefore, Fu as well as Mahyar are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Fu with the teachings of Mahyar by training each of the models before model selection and using bias as model parameter. Mahyar provides as additional motivation for combination that each model may have distinct benefits ([Col. 7 l. 30-45] “The different subsets are then used as different training data for the various voice synthesis models, such that a trained voice synthesis model is specific to a particular speaker (e.g., in contrast to a trained voice synthesis model learning to replicate multiple speakers). It should be appreciated that the same voice synthesis model structure (e.g., the hyper parameters for a particular neural network) may be in common for modeling speaker A and speaker B, but each speaker would be associated with a separate training process such that the same neural network structure would have different weight/biases, etc., following the training process.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 16, claim 16 is directed towards the method performed by the apparatus of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 16. Regarding claim 17, claim 17 is substantially similar to claim 1. Therefore, the rejection applied to claim 1 also applies to claim 17. Claims 2-5, 7-11, and 15 are rejected under U.S.C. §103 as being unpatentable over the combination of Fu and Mahyar and in further view of Koch (US20180240041A1). Regarding claim 2, the combination of Fu and Mahyar teaches The learning apparatus according to claim 1. However, the combination of Fu and Mahyar doesn't explicitly teach further comprising: generate a plurality of input values to be input to a predetermined first function that calculates a random number value based on the input value, and generates, for each of the generated input values, a plurality of models having parameters corresponding to the random number values output from the predetermined first function when the input values have been input. Koch, in the same field of endeavor, teaches generate a plurality of input values to be input to a predetermined first function that calculates a random number value based on the input value, and generates, for each of the generated input values, ([¶0131] "a random seed value may be specified for each search method that may be the same for all search methods or may be defined separately for each search method") a plurality of models having parameters corresponding to the random number values output from the predetermined first function when the input values have been input.([¶0134] "the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters. If the Random search method is selected, a sample size value may be specified for all or for each hyperparameter that defines the number of hyperparameter configurations to evaluate in a single iteration."). The combination of Fu and Mahyar as well as Koch are directed towards neural architecture search. Therefore, the combination of Fu and Mahyar as well as Koch are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fu and Mahyar with the teachings of Koch by using the model generation method disclosed in Koch for the model randomization in Fu. Koch provides that a variety of search methods may be used, each having distinct advantages, for example Koch provides as additional motivation for combination in using the LHS method ([¶0135] “ensures that the lower and upper bounds of the hyperparameter tuning range are included, and for discrete hyperparameters with a number of levels less than the requested sample size”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 3, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 2, wherein the plurality of input values comprises input values with random number values that satisfy a predetermined condition.(Koch [¶0134] "the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters" [¶0135] "the LHS search method generates uniform hyperparameter values across the range of each hyperparameter and randomly combines them across hyperparameters. If the hyperparameter is continuous or discrete with more levels than a requested sample size, a uniform set of samples is taken across the hyperparameter range including a lower and an upper bound" Satisfying upper and lower bounds interpreted as predetermined condition). Regarding claim 4, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 3, wherein the plurality of input values comprises input values with random number value that fall within a predetermined range.(Koch [¶0134] "the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters"). Regarding claim 5, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 3, wherein the plurality of input values comprises input values with a distribution of associated random number values having a predetermined probability distribution. (Koch [¶0135] "the LHS search method generates uniform hyperparameter values across the range of each hyperparameter and randomly combines them across hyperparameters. If the hyperparameter is continuous or discrete with more levels than a requested sample size, a uniform set of samples is taken across the hyperparameter range including a lower and an upper bound" [¶0133] " the Grid search method generates uniform hyperparameter values across the range of each hyperparameter and combines them across hyperparameters" Koch explicitly states that the generated random number values have uniform probability distribution). Regarding claim 7, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 2, wherein a function in which the distribution of the random number values output when the input value has been input indicates a predetermined probability distribution is selected as the predetermined first function(Koch [¶0135] "the LHS search method generates uniform hyperparameter values across the range of each hyperparameter and randomly combines them across hyperparameters. If the hyperparameter is continuous or discrete with more levels than a requested sample size, a uniform set of samples is taken across the hyperparameter range including a lower and an upper bound" [¶0133] " the Grid search method generates uniform hyperparameter values across the range of each hyperparameter and combines them across hyperparameters" Koch explicitly states that the generated random number values have uniform probability distribution where uniform probability distribution is explicitly selected as the first function). Regarding claim 8, the combination of Fu and Mahyar teaches The learning apparatus according to claim 1. However, the combination of Fu and Mahyar doesn't explicitly teach wherein the selecting of the one of the plurality of models based on model accuracy comprises selecting a model that satisfies predetermined evaluation conditions from among the plurality of trained models. Koch, in the same field of endeavor, teaches the selecting of the one of the plurality of models based on model accuracy comprises selecting a model that satisfies predetermined evaluation conditions from among the plurality of trained models([¶0137] "A tournament selection process may be used to randomly choose a group of members from the current population, compare their fitness, and select the fittest from the group to propagate to the next generation" [¶0138] "growth steps may be performed each iteration to permit selected hyperparameter configurations of the population (based on diversity and fitness) to benefit from local optimization over the continuous variables" [¶0240] "Referring to FIG. 17, a model improvement (error reduction or accuracy increase where higher is better) for the suite of ten common machine learning test problems illustrated in FIG. 16 is shown" Examiner notes that both the training and tournament selection tuning process satisfy this claim). The combination of Fu and Mahyar as well as Koch are directed towards neural architecture search. Therefore, the combination of Fu and Mahyar as well as Koch are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fu and Mahyar with the teachings of Koch by using the model generation method disclosed in Koch for the model randomization in Fu. Koch provides that a variety of search methods may be used, each having distinct advantages, for example Koch provides as additional motivation for combination in using the LHS method ([¶0135] “ensures that the lower and upper bounds of the hyperparameter tuning range are included, and for discrete hyperparameters with a number of levels less than the requested sample size”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 9, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 8, wherein the predetermined evaluation conditions comprise a change in the evaluation value satisfies a predetermined mode.(Koch [¶0137] "A tournament selection process may be used to randomly choose a group of members from the current population, compare their fitness, and select the fittest from the group to propagate to the next generation" [¶0138] "growth steps may be performed each iteration to permit selected hyperparameter configurations of the population (based on diversity and fitness) to benefit from local optimization over the continuous variables" [¶0240] "Referring to FIG. 17, a model improvement (error reduction or accuracy increase where higher is better) for the suite of ten common machine learning test problems illustrated in FIG. 16 is shown" Examiner notes that both the training and tournament selection tuning process satisfy this claim). Regarding claim 10, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 9, wherein model selection is based on the change in the evaluation value during iterative learning of the features of a part of the predetermined learning data a predetermined number of times satisfies the predetermined mode.(Koch [¶0137] "a maximum number of iterations may be specified where the population size defines the number of hyperparameter configurations to evaluate each iteration"). Regarding claim 11, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 8, wherein the predetermined evaluation conditions comprise a plurality of conditions designated by the user(Koch [¶0148] "As another option, the user can select a hyperparameter configuration included in the “Tuner Results” output table that is less complex, but provides a similar objective function value in comparison to the hyperparameters included in the “Best Configuration” output table"). Regarding claim 15, the combination of Fu and Mahyar teaches The learning apparatus according to claim 1. However, the combination of Fu and Mahyar doesn't explicitly teach wherein the selecting of one of the plurality of models based on model accuracy comprises selecting one of the models based on model accuracy for each model having different parameters. Koch, in the same field of endeavor, teaches the selecting of one of the plurality of models based on model accuracy comprises selecting one of the models based on model accuracy for each model having different parameters ([Abstract] "scoring of the trained model using a validation dataset and the assigned hyperparameter configuration is requested to compute an objective function value, and the received objective function value and the assigned hyperparameter configuration are stored. A best hyperparameter configuration is identified based on an extreme value of the stored objective function values" [¶0101] "F1 uses an F1 coefficient as the objective function [...] MCE uses a misclassification rate as the objective function" [¶0239] "Referring to FIG. 16, a final tuned model error—as averaged across ten tuning runs that used different validation partitions—for each problem and each modeling algorithm is shown for a suite of ten common machine learning test problems" [¶0240] "Referring to FIG. 17, a model improvement (error reduction or accuracy increase where higher is better) for the suite of ten common machine learning test problems illustrated in FIG. 16 is shown" Koch explicitly and repeatedly evaluates many hyperparameter configurations (plurality of models having different parameters), selecting the best based on an extreme objective value (explicitly anticipated as being accuracy)). The combination of Fu and Mahyar as well as Koch are directed towards neural architecture search. Therefore, the combination of Fu and Mahyar as well as Koch are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fu and Mahyar with the teachings of Koch by using the model generation method disclosed in Koch for the model randomization in Fu. Koch provides that a variety of search methods may be used, each having distinct advantages, for example Koch provides as additional motivation for combination in using the LHS method ([¶0135] “ensures that the lower and upper bounds of the hyperparameter tuning range are included, and for discrete hyperparameters with a number of levels less than the requested sample size”). This motivation for combination also applies to the remaining claims which depend on this combination. Claim 6 is rejected under U.S.C. §103 as being unpatentable over the combination of Fu and Mahyar and Koch and in further view of Detwiler (US20170329875A1). Regarding claim 6, the combination of Fu, Mahyar, and Koch teaches The learning apparatus according to claim 3. However, the combination of Fu, Mahyar, and Koch doesn't explicitly teach, wherein the plurality of input values comprises input values with a mean value of associated random number values meeting a predetermined value. Detwiler, in the same field of endeavor, teaches the plurality of input values comprises input values with a mean value of associated random number values meeting a predetermined value.([0157] "Here, N(0,1) is a random number according to a normal distribution with mean zero and expectation one"). The combination of Fu, Mahyar, and Koch as well as Detwiler are directed towards machine learning systems with random number generation. Therefore, The combination of Fu, Mahyar, and Koch as well as Detwiler are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of The combination of Fu, Mahyar, and Koch with the teachings of Detwiler by substituting a normal distribution for the uniformly distributed random number generator used in Koch. While it would have been obvious to one of ordinary skill in the art that the uniform distribution used in Koch could have a mean of zero and that a standard normal distribution has a mean of zero, Detwiler explicitly teaches that using a normal distribution with a mean of zero over the same exact range used in Koch in an analogous evolutionary algorithm is known and would lead to predictable results. Claims 12, 13, and 14 are rejected under U.S.C. §103 as being unpatentable over the combination of Fu and Mahyar and Parnell (US 20200184369 A1). Regarding claim 12, the combination of Fu and Mahyar teaches The learning apparatus according to claim 1. However, the combination of Fu and Mahyar doesn't explicitly teach wherein the processor is further configured to generate a plurality of input values to be input to a predetermined second function that calculates a random number value for each input value; generate, for each of the plurality of input values, a part of the predetermined learning data based on corresponding random number value output by the predetermined second function, wherein the part of the predetermined learning data is used to train each of the plurality of models. Parnell, in the same field of endeavor, teaches the processor is further configured to generate a plurality of input values to be input to a predetermined second function that calculates a random number value for each input value; generate, for each of the plurality of input values, a part of the predetermined learning data based on corresponding random number value output by the predetermined second function, wherein the part of the predetermined learning data is used to train each of the plurality of models. ([¶0007] "For successive batches of the training data, defined by respective subsets of one of the row coordinates and column coordinates, the method includes generating, in the host computer, random numbers associated with respective coordinates in a current batch b and sending the random numbers to the accelerator unit. In parallel with generating the random numbers for batch b, batch b is copied from the host computer to the accelerator unit. The method further comprises, in the accelerator unit and in parallel with the copying of batch b, sorting the random numbers for coordinates in the previous batch (b−1) to randomly permute the coordinates and performing the stochastic optimization process for the permuted coordinates in batch (b−1) to update the model vector w in dependence on coordinates in that batch."). The combination of Fu and Mahyar as well as Parnell are directed towards machine learning training. Therefore, the combination of Fu and Mahyar as well as Parnell are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fu and Mahyar with the teachings of Parnell by using a random number generator for the training dataset order. Parnell provides as additional motivation for combination ([¶0008] “The random numbers can be generated efficiently on the host computer and then sent to the accelerator ready for processing of the next batch of data. The random numbers are then sorted on the accelerator, where the sorting operation can be performed with high efficiency, to obtain the required coordinate permutation. Performing the tasks in parallel in this way provides more effective use of system resources, offering significant improvement in efficiency of machine learning operations.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 13, the combination of Fu, Mahyar, and Parnell teaches The learning apparatus according to claim 12, wherein input value generation is iteratively performed to generate learning data as a learning target in the learning for iterative model training(Parnell [¶0007] "For successive batches of the training data, defined by respective subsets of one of the row coordinates and column coordinates, the method includes generating, in the host computer, random numbers associated with respective coordinates in a current batch b and sending the random numbers to the accelerator unit. In parallel with generating the random numbers for batch b, batch b is copied from the host computer to the accelerator unit. The method further comprises, in the accelerator unit and in parallel with the copying of batch b, sorting the random numbers for coordinates in the previous batch (b−1) to randomly permute the coordinates and performing the stochastic optimization process for the permuted coordinates in batch (b−1) to update the model vector w in dependence on coordinates in that batch." Batch interpreted as a time of repeated learning.). Regarding claim 14, the combination of Fu, Mahyar, and Parnell teaches The learning apparatus according to claim 12, further comprising: generating, as a part of the predetermined learning data, learning data in which the random number values are associated as a learning order.(Parnell [¶0007] "For successive batches of the training data, defined by respective subsets of one of the row coordinates and column coordinates, the method includes generating, in the host computer, random numbers associated with respective coordinates in a current batch b and sending the random numbers to the accelerator unit. In parallel with generating the random numbers for batch b, batch b is copied from the host computer to the accelerator unit. The method further comprises, in the accelerator unit and in parallel with the copying of batch b, sorting the random numbers for coordinates in the previous batch (b−1) to randomly permute the coordinates and performing the stochastic optimization process for the permuted coordinates in batch (b−1) to update the model vector w in dependence on coordinates in that batch."). Claim 18 is rejected under U.S.C. §103 as being unpatentable over the combination of Koch and Fu. Regarding claim 18, Koch teaches A learning apparatus comprising: a processor,([¶0006] " The computing device includes, but is not limited to, a processor and a non-transitory computer-readable medium operably coupled to the processor. The computer-readable medium has instructions stored thereon that, when executed by the processor, cause the computing device to automatically select hyperparameter values based on objective criteria for training a predictive model.") wherein the processor is configured to: generate a plurality of random number seeds and input the plurality of random number seeds into a predetermined function to generate a plurality of random number values that correspond to the plurality of random number seeds; ([¶0131] "a random seed value may be specified for each search method that may be the same for all search methods or may be defined separately for each search method" [¶0134] "the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters." Koch explicitly generates plural random seed values used for/as input for a search method that generates a plurality of random hyperparameter values) generate, for the plurality of random number seeds, a plurality of models with differing parameters based on the plurality of random number values;([¶0134] "the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters." Koch explicitly generates plural random seed values used for/as input for a search method that generates a plurality of random hyperparameter values, the hyperparameter values determining model configurations) train the plurality of models to learn features of a part of predetermined learning data; ([¶0005] "cause the computing device to automatically select hyperparameter values based on objective criteria for training a predictive model. A plurality of tuning evaluation parameters that include a model type, a search method type, and values to evaluate for each hyperparameter of a plurality of hyperparameters associated with the model type are accessed. A number of session computing devices allocated to each session of a plurality of sessions is determined. Each session computing device of the number of session computing devices processes a subset of an input dataset. [...] The model is trained using the assigned hyperparameter configuration and a training dataset that is a first portion of the input dataset") select one of the plurality of models based on model accuracy; ([¶0100] "The objective function specifies a measure of model error (performance) to be used to identify a best configuration of the hyperparameters among those evaluated" [¶0101] "F1 uses an F1 coefficient as the objective function [...] MCE uses a misclassification rate as the objective function" [¶0239] "Referring to FIG. 16, a final tuned model error—as averaged across ten tuning runs that used different validation partitions—for each problem and each modeling algorithm is shown for a suite of ten common machine learning test problems" [¶0240] "Referring to FIG. 17, a model improvement (error reduction or accuracy increase where higher is better) for the suite of ten common machine learning test problems illustrated in FIG. 16 is shown" Koch explicitly discloses that the model tuning algorithm selects a best model from a plurality of models to improve (tune) a base models error/accuracy) and train the selected model to learn features of the predetermined learning data, ([¶0188] "a final hyperparameter configuration is selected based on the hyperparameter configuration that generated the best or lowest objective function value" [¶0191] "the selected session is requested to execute the final hyperparameter configuration based on the parameter values in the data structure. In an illustrative embodiment, a train request is sent to session manager device 400 of the selected session to execute the “train” action based on the selected model type. [...] Characteristics that define the trained model using the final hyperparameter configuration are provided back to the main thread on which selection manager application 312 is instantiated" Koch explicitly performs a post-selection "final model" training using the best configuration and stores that trained selected model for later use.) wherein the plurality of random number values satisfies a predetermined condition.([¶0134] " the Random search method randomly generates hyperparameter values across the range of each hyperparameter and combines them across hyperparameters" [¶0135] "the hyperparameter range including a lower and an upper bound" The range of hyperparameter values interpreted as satisfying a predetermined condition (upper and lower bound)). However, Koch does not explicitly teach wherein a prediction process using the trained selected model includes a plurality of processes, and the learning apparatus decides, based on features of the trained selected model, which arithmetic unit among a plurality of arithmetic units having different architectures is to execute each of the plurality of processes included in the prediction process, to reduce processing time for the prediction process. Fu, in the same field of endeavor, teaches wherein a prediction process using the trained selected model includes a plurality of processes, and the learning apparatus decides, based on features of the trained selected model, which arithmetic unit among a plurality of arithmetic units having different architectures is to execute each of the plurality of processes included in the prediction process, to reduce processing time for the prediction process. ([p. 47] "NAS techniques 4–6 target to identify efficient models on single CPU/GPU systems" [p. 50] "our expert placement policy uniformly distributes each concatenation node and the associated operations onto all existing devices [...] considering the heterogeneous property of the underlying computing platforms. To further optimize the latency of the searched architecture G, the last step of the ColocNAS leverages policy gradient for device placement on the given hardware"). Koch as well as Fu are directed towards neural architecture search. Therefore, Koch as well as Fu are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Koch with the teachings of Fu by using the model architecture search in Koch for device placement. Fu provides as additional motivation for combination that the reinforcement learning device placement and fine tuning can optimize the model ([p. 48] “to reduce its runtime on the target platform.”). This motivation for combination also applies to the remaining claims which depend on this combination. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Show 4 earlier events
Dec 03, 2025
Response after Non-Final Action
Jan 05, 2026
Request for Continued Examination
Jan 07, 2026
Response after Non-Final Action
Feb 13, 2026
Non-Final Rejection mailed — §103, §112
May 15, 2026
Examiner Interview Summary
May 15, 2026
Applicant Interview (Telephonic)
Jun 04, 2026
Response Filed
Aug 18, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month