Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The amendments filed on 06/30/2026 have been considered. Claims 1, 11, 12, 17, 19, 20 have been amended. Claims 7-10 are cancelled. Thus, claims 1-6, 11-20 are pending and presented for
examination.
Applicant's amendment filed on 06/30/2026 with respect to the title objection have been fully considered. Thus, the title objection is withdrawn.
Applicant's amendment filed on 06/30/2026 with respect to the 35 U.S.C. 112(f) claim interpretation have been fully considered. Thus, the 35 U.S.C. 112(f) claim interpretation is withdrawn.
Applicant's amendments filed on 06/30/2026 with respect to the 35 U.S.C. 112(b) rejection have been fully considered. Thus, the 35 U.S.C. 112(b) rejection is withdrawn.
Applicant's arguments filed on 06/30/2026 with respect to the 35 U.S.C. 101 rejections have
been fully considered and finds them persuasive, Thus, the 35 U.S.C. 101 rejections have been withdrawn.
Applicant's arguments filed on 06/30/2026 with respect to the 35 U.S.C. 103 rejections have
been fully considered but are moot because of the new ground of rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4-5, and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Jacobs et al. (“Adaptive Mixtures of Local Experts”, hereinafter “Jacobs”) in view of patent application US 2018/0137427 A1 Hsieh et al., hereinafter “Hsieh” further in view of non-patent literature XGboost (“xgboost Release 1.0.2”, hereinafter “XGboost”).
Claim 1
Jacobs teaches:
An information processing …, (page 79, “We present a new supervised learning procedure for systems composed of many separate networks…”) comprising: … an ensemble learning-type inference model that performs inference based on each inference result by a plurality of inference models; (Page 80, “the final output of the whole system is a linear combination of the outputs of the local experts, with the gating network determining the proportion of each local output in the linear combination” – EN: this denotes an ensemble type model that generates a final system output by blending or combining the individual predictions of multiple local expert networks.) … input data and correct answer data that corresponds to the input data; (Page 80, “dc is the desired output vector in case c.” – EN: this denotes case c which is the training case used to train the networks and dc is the desired output vector in case c.)
… inferred output data of the ensemble learning-type inference model by inputting the input data to the ensemble learning-type inference model; and (Page 80, “oic is the output vector of expert i on case c, pic is the proportional contribution of expert i to the combined output vector” Page 81, “Figure 1: A system of expert and gating networks. Each expert is a feedforward network and all experts receive the same input and have the same number of outputs.” – EN: this denotes generating the inferred output data (the combined output vector) of the ensemble model by inputting the input data of case c to the expert networks that constitute the ensemble.)
… with respect to a part of or all of each of the inference models that constitute the ensemble learning-type inference model by using an update amount based on the inferred output data and the correct answer data,
PNG
media_image1.png
111
1046
media_image1.png
Greyscale
(page 80, “This error measure compares the desired output with a blend of the outputs of the local experts, so, to minimize the error, each local expert must make its output cancel the residual error that is left by the combined effects of all the other experts. When the weights in one expert change, the residual error changes, and so the error derivatives for all the other local experts change.” – EN: this denotes calculating a global error by taking the difference between the target answer and the combined ensemble output. It further denotes updating the weights of the individual expert models using error derivatives (the update amount) that are directly based on minimizing this global residual error.)
Jacobs does not explicitly disclose:
“apparatus”, “a memory that stores”, “data acquiring processor circuitry configured to acquire”, “inferred output data generating processor circuitry configured to generate”, “an additional learning processor configured to perform additional learning processing”, “and to store the updated ensemble learning-type inference model in the memory”
However, Hsieh teaches:
“apparatus” (Para 2, “The disclosure relates in general to an ensemble learning prediction apparatus”)
“a memory that stores” (Para 21, “The adaptive ensemble weight is stored in a data medium.” Para 38, “In some embodiments, the non-transitory computer-readable storage medium can be stored in a computer program product having instructions allocated to a computing device executing the abovementioned method.” – EN: the data medium that stores the ensemble weight of the ensemble learning prediction and the storage medium of the computing device constitute a memory that stores the ensemble learning-type inference model.)
“data acquiring processor circuitry configured to acquire” (Para 6, “According to one embodiment, an ensemble learning prediction apparatus is provided. The apparatus comprises a loss module receiving a sample data” – Examiner’s Note (EN): FIG. 2 also illustrates that these modules reside and operate within a processor.)
“inferred output data generating processor circuitry configured to generate” (FIG. 2, EN: Fig. 2 illustrates the generation of a prediction “output” from the weighting module of the ensemble learning prediction apparatus)
“an additional learning processor configured to perform additional learning processing” (Para 3, “However, in practical application, as the environment varies with the time, concept drifting phenomenon may occur, and the accuracy of the ensemble learning model created according to historical data will decrease. Under such circumstances, the prediction model must be re-trained or adjusted by use of newly created data to restore the prediction accuracy within a short period of time” Para 21, “The adaptive ensemble weight can be adapted to the current environment to resolve the concept drifting problem.”)
“and to store the updated ensemble learning-type inference model in the memory” (Para 21, “The ensemble learning prediction apparatus of the present disclosure obtains an adaptive ensemble weight by adjusting the ensemble weight by use of historical ensemble weights. The adaptive ensemble weight is stored in a data medium.” – EN: the adjusted (updated) ensemble weight of the model is stored in the data medium (memory), which constitutes storing the updated ensemble learning-type inference model in the memory.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the teachings of Jacobs with the teachings of Hsieh. Specifically, it would have been obvious to integrate Jacobs’ method of updating the internal parameters of individual base models using an error measure based on the combined ensemble output and the correct answer with Hsieh’s hardware and ensemble learning apparatus configured to perform continuous, additional learning. The motivation for doing so would be to enable the individual base models (the local experts) to continuously adapt in response to changing environments (concept drift). See Para 3 in Hsieh, “However, in practical application, as the environment varies with the time, concept drifting phenomenon may occur, and the accuracy of the ensemble learning model created according to historical data will decrease. Under such circumstances, the prediction model must be re-trained or adjusted by use of newly created data to restore the prediction accuracy within a short period of time…”
Jacobs in view of Hsieh does not explicitly disclose:
wherein the inference models are trained decision trees, and
wherein the additional learning processor performs the additional learning processing by adding the update amount to an output value corresponding to a terminal node of each of the trained decision trees, whereby changes related to the structure of the trained decision trees do not occur by the additional learning processing and an additional storage capacity of the memory is not required.
However, XGboost teaches:
wherein the inference models are trained decision trees, and (Page 15, “The tree ensemble model consists of a set of classification and regression trees (CART).” Page 15, “What is actually used is the ensemble model, which sums the prediction of multiple trees together.” Page 59, “update: Starts from an existing model and only updates its trees.” - EN: the existing model of the update process is a previously trained model, and its trees are trained decision trees. The trees whose predictions are summed to form the ensemble output constitute the plurality of inference models, corresponding to the local experts whose outputs are combined in Jacobs.)
wherein the additional learning processor performs the additional learning processing by adding (...) [the update amount] to an output value corresponding to a terminal node of each of the trained decision trees, (Page 15, “We classify the members of a family into different leaves, and assign them the score on the corresponding leaf… In CART, a real score is associated with each of the leaves”. Page 58, “refresh: refreshes tree’s statistics and/or leaf values based on the current data.” Page 59, “This is a parameter of the refresh updater. When this flag is 1, tree leafs as well as tree nodes’ stats are updated.” Page 59, “In each boosting iteration, a tree from the initial model is taken, a specified sequence of updaters is run for that tree, and a modified tree is added to the new model.” - EN: a leaf of a decision tree is a terminal node, and the real score associated with the leaf is the output value corresponding to the terminal node. Running the update process with the refresh updater and refresh_leaf=1 updates the leaf values of each tree of the trained model based on the current data. Thus, XGboost teaches adding an updating value to the output value corresponding to a terminal node of each of the trained decision trees. XGboost is not relied upon for the particular "update amount" recited; the update amount, which is based on the inferred output data and the correct answer data, is taught by Jacobs as set forth above (Jacobs, Page 80, Eq. 1.1). In the combination, the update amount of Jacobs is the value by which the leaf values of XGboost's trained decision trees are updated. See the motivation for the combination below.)
whereby changes related to the structure of the trained decision trees do not occur by the additional learning processing and an additional storage capacity of the memory is not required. (Page 59, “With process_type=update, one cannot use updaters that create new trees.” Page 59, “The new model would have either the same or smaller number of trees, depending on the number of boosting iterations performed.” - EN: the refresh updater updates only the node statistics and leaf values of the existing trees; the updaters that construct or prune tree structure are separate updaters that are not run (Page 58), and no node of any tree is created, removed, or re-split, whereby changes related to the structure of the trained decision trees do not occur by the additional learning processing. Because no new tree or node is created and the update overwrites existing leaf values in place, the updated model occupies no more memory than the trained model it replaces, whereby an additional storage capacity of the memory is not required.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble additional learning of Jacobs and Hsieh, in which each inference model is updated using an update amount based on the inferred output data and the correct answer data, with the trained decision trees and the leaf-value update process of XGboost, such that the update amount of Jacobs is the value added to the output values corresponding to the terminal nodes (leaves) of XGboost's trained decision trees when the additional learning is performed. The motivation for doing so would be to perform the re-training on newly created data taught by Hsieh (Hsieh, Para 3) with a small computation cost and without increasing the size of the stored model, because updating only the leaf values of the existing trees avoids constructing new trees. As XGboost elaborates regarding the update process type on page 59, "update: Starts from an existing model and only updates its trees... The new model would have either the same or smaller number of trees, depending on the number of boosting iterations performed."
Claim 4
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Jacobs further teaches:
wherein the update amount is a value based on a difference between the inferred output data and the correct answer data. (Page 80 equation 1.1)
PNG
media_image2.png
183
1046
media_image2.png
Greyscale
Claim 5
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Jacobs further teaches:
wherein the update amount is a value based on a value obtained by multiplying a difference between the inferred output data and the correct answer data by a learning rate. (Page 80 equation 1.1 and Page 85, “All simulations were performed using a simple gradient descent algorithm with fixed step size t” – EN: this denotes using a fixed step size in a gradient descent algorithm which is synonymous with a learning rate. When a gradient descent algorithm is applied to the error function in equation 1.1, the mathematical derivative extracts the difference term. The algorithm then multiplies this gradient by the step size to calculate the final weight.)
Claim 16
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Hsieh further teaches:
wherein the additional learning processing is online learning. (Para 17, “To avoid the online learning sample data being over-trained and losing the required diversity between the basic hypotheses of the ensemble learning model, the diversity of hypotheses is considered during the learning process.” Para 21, “{x.sub.n.sup.(t)}, n=1,2, . . . N denotes an online learning sample data in the t-th block”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs to perform the additional learning as online learning as taught by Hsieh. The motivation for doing so would be to enable the individual base models (the local experts) to continuously adapt in response to changing environments (concept drift). See Para 3 in Hsieh, “However, in practical application, as the environment varies with the time, concept drifting phenomenon may occur, and the accuracy of the ensemble learning model created according to historical data will decrease. Under such circumstances, the prediction model must be re-trained or adjusted by use of newly created data to restore the prediction accuracy within a short period of time…”
Claim 17
Jacobs teaches:
An information processing method ... (Page 79, "We present a new supervised learning procedure for systems composed of many separate networks, each of which learns to handle a subset of the complete set of training cases." -- EN: the supervised learning procedure constitutes an information processing method.) ... an ensemble learning-type inference model that performs inference based on each inference result by a plurality of inference models, comprising: (Page 80, "the final output of the whole system is a linear combination of the outputs of the local experts, with the gating network determining the proportion of each local output in the linear combination")
Jacobs does not explicitly disclose the method being "performed on an information processing apparatus having a memory that stores" the model. However, Hsieh teaches an apparatus with a memory that stores the model (Para 2; Para 21, "The adaptive ensemble weight is stored in a data medium."), as set forth in the rejection of claim 1.
The remaining limitations of claim 17 are substantially the same as claim 1, therefore claim 17 is rejected under the same rationale as claim 1.
Claim 18
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 17, Hsieh further teaches:
A non-transitory computer-readable medium (Para 8, “The non-transitory computer-readable storage medium provided in the present disclosure can execute the abovementioned method.”) having one or more executable instructions stored thereon (Para 38, “stored in a computer program product having instructions allocated to a computing device”) causing a computer to function as an information processing device which, (Para 38, “allocated to a computing device”, Para 6, “According to one embodiment, an ensemble learning prediction apparatus is provided.”) when executed by processor circuitry, cause the processor circuitry to perform the information processing method (Para 38, “allocated to a computing device executing the abovementioned method.” Para 8, “The non-transitory computer-readable storage medium provided in the present disclosure can execute the abovementioned method.”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine ensemble learning network of Jacobs with the non-transitory computer-readable medium of Hsieh. The motivation for doing so would be to allow the ensemble learning method to be implemented on a functional component. See para 38 of Hsieh, “In some embodiments, the non-transitory computer-readable storage medium can be stored in a computer program product having instructions allocated to a computing device executing the abovementioned method.”
Claim 19
Hsieh teaches:
An information processing system (Para 2, “The disclosure relates in general to an ensemble learning prediction apparatus”)
The remaining limitations of claim 19 are substantially the same as claim 1, therefore claim 19 is rejected under the same rationale as claim 1.
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine ensemble learning network of Jacobs with the apparatus of Hsieh. The motivation for doing so would be to allow the ensemble learning method to be implemented on a functional component. See para 6 in Hsieh, “According to one embodiment, an ensemble learning prediction apparatus is provided. The apparatus comprises a loss module receiving a sample data and calculating a loss according to a first prediction result of the sample data and an actual result…”
Claims 2-3 and 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Jacobs et al. (“Adaptive Mixtures of Local Experts”, hereinafter “Jacobs”) in view of patent application US 2018/0137427 A1 Hsieh et al., hereinafter “Hsieh” in view of non-patent literature XGboost (“xgboost Release 1.0.2”, hereinafter “XGboost”) further in view of Hastie et al. (“The Elements of Statistical Learning Data Mining, Inference, and Prediction”, hereinafter “Hastie”)
Claim 2
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Hastie teaches:
wherein the ensemble learning-type inference model is a boosting learning-type inference model which is constituted by a plurality of inference models formed by sequential learning (Page 338, “The purpose of boosting is to sequentially apply the weak classification algorithm to repeatedly modified versions of the data, thereby producing a sequence of weak classifiers Gm(x), m = 1, 2, . . . , M.”) so that each of the inference models reduces an inference error due to a higher-order inference model group. (Page 338-339, “Each successive classifier is thereby forced to concentrate on those training observations that are missed by previous ones in the sequence.”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with boosting model of Hastie. The motivation for doing so would be to improve the overall predictive accuracy of the model by combining multiple weak learners to create an accurate and powerful “committee” of classifiers. See page 337 in Hastie that states, “The motivation for boosting was a procedure that combines the outputs of many “weak” classifiers to produce a powerful “committee.”
Claim 3
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Hastie teaches:
wherein the ensemble learning-type inference model is a bagging learning-type inference model (Page 282, “Bootstrap aggregation or bagging averages this prediction over a collection of bootstrap samples”) which performs inference based on each inference result of a plurality of inference models, (Page 282, “For each bootstrap sample Z∗b, b = 1, 2, . . . , B, we fit our model, giving prediction ˆf∗b(x).”) each formed by learning based on a plurality of data groups extracted from a same learning target data group. (Page 282, “Bootstrap aggregation or bagging averages this prediction over a collection of bootstrap samples… The bagging estimate is defined by… (equation 8.51)”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the bagging model of Hastie. The motivation for doing so would be to improve the model’s accuracy by reducing its variance. See page 282 in Hastie that states, “Here we show how to use the bootstrap to improve the estimate or prediction itself… Bootstrap aggregation or bagging averages this prediction over a collection of bootstrap samples, thereby reducing its variance.” And Page 283, “Bagging can dramatically reduce the variance of unstable procedures like trees, leading to improved prediction.”
Claim 11
Jacobs in view of Hsieh in view of XGboost further in view of Hastie teaches all the limitations of claim 2, Hsieh further teaches:
first output data generating processor circuitry configured to generate (FIG. 2, EN: Fig. 2 illustrates the generation of a prediction “output” from the weighting module of the ensemble learning prediction apparatus)
second output data generating processor circuitry configured to generate (FIG. 2, EN: Fig. 2 illustrates the generation of a prediction “output” from the weighting module of the ensemble learning prediction apparatus)
final output data generating processor circuitry configured to generate (FIG. 2, EN: Fig. 2 illustrates the generation of a prediction “output” from the weighting module of the ensemble learning prediction apparatus)
…additional learning processing is processing of… (Para 3, “However, in practical application, as the environment varies with the time, concept drifting phenomenon may occur, and the accuracy of the ensemble learning model created according to historical data will decrease. Under such circumstances, the prediction model must be re-trained or adjusted by use of newly created data to restore the prediction accuracy within a short period of time” Para 21, “The adaptive ensemble weight can be adapted to the current environment to resolve the concept drifting problem.”)
Jacobs in view of Hsieh in view of XGboost does not explicitly disclose:
wherein the boosting learning-type inference model further includes a first inference model, the first inference model comprising: a first output … first output data by inputting the input data to a first approximate function generated based on training input data and training correct answer data that corresponds to the training input data; a second output … second output data by inputting the input data to a second trained model generated by performing machine learning based on the training input data and difference data between output data generated by inputting the training input data to the first approximate function and the training correct answer data; and a final output … final output data based on the first output data and the second output data, wherein the … updating the second trained model using an update amount based on difference data between the correct answer data and the first output data and the inferred output data.
However, Hastie teaches:
wherein the boosting learning-type inference model further includes a first inference model, the first inference model comprising: a first output … first output data by inputting the input data to a first approximate function generated based on training input data and training correct answer data that corresponds to the training input data; (Page 342, “Forward stagewise modeling approximates the solution to (10.4) by sequentially adding new basis functions to the expansion without adjusting the parameters and coefficients of those that have already been added.” – EN: The “already added” portion of the expansion which is mathematically represented as fm−1(x) in the text acts as the first output or rough approximation that remains unadjusted during the current step.) a second output … second output data by inputting the input data to a second trained model generated by performing machine learning based on the training input data and difference data between output data generated by inputting the training input data to the first approximate function and the training correct answer data; (Page 343, “Thus, for squared-error loss, the term βmb(x; ym) that best fits the current residuals is added to the expansion at each step.” And Page 342 - Algorithm 10.2 Forward Stagewise Additive Modeling, specifically step (b). – EN: The term βmb(x; ym) is the newly trained model (second output) designed to fit the residuals (errors). Algorithm 10.2 shows these two parts are mathematically combined by adding for the final result fm(x).) and a final output … final output data based on the first output data and the second output data, wherein the … updating the second trained model using an update amount based on difference data between the correct answer data and the first output data and the inferred output data. (Page 343, “where rim = yi − fm−1(xi) is simply the residual of the current model on the ith observation.” Page 342, “At each iteration m, one solves for the optimal basis function b(x; ym) and corresponding coefficient βm to add to the current expansion fm−1(x).” – EN: this denotes calculating the difference (residual rim) between the correct answer (yi) and the first output (fm−1(x)), and uses that difference to solve for the new second model.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the forward stagewise additive modeling of Hastie. The motivation for doing so would be to dynamically correct for errors adapt to changing environments (such as concept drift) by training a new model specifically on the residual error s of a fixed baseline model, thereby adapting to new data without disrupting or having to retrain the already-established baseline accuracy. Hastie details this architecture of freezing prior models and fitting new ones to the residuals in page 342 stating that “forward stagewise modeling approximates the solution to (10.4) by sequentially adding new basis functions to the expansion without adjusting the parameters and coefficients of those that have already been added.” Hastie further explains that at each iteration, the new model is fit to the error of the frozen baseline in page 343 stating that “where rim = yi − fm−1(xi) is simply the residual of the current model on the ith observation”.
Claim 12
Jacobs in view of Hsieh in view of XGboost further in view of Hastie teaches all the limitations of claim 11, Hastie further teaches:
wherein in the boosting learning-type inference model, only inference models equal to or lower than a predetermined inference model are configured as the first inference model. (Page 361 - Algorithm 10.3 Gradient Tree Boosting Algorithm – EN: this denotes algorithm 10.3 which initializes at step 1 with a basic starting point without the two-part structure. The hybrid two-part structure only begins at position m = 1 and applies to all subsequent models lower in the chain.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the gradient tree boosting initialization architecture of Hastie. The motivation for doing so would be to provide a structured mathematical approach to the teaching of Jacobs and Hsieh by using an initialization step as a frozen baseline model that captures the initial global data distribution, allowing the system to handle concept drift by only calculating and updating the residuals in the subsequent model layers. See Algorithm 10.3 on page 361.
Claim 13
Jacobs in view of Hsieh in view of XGboost further in view of Hastie teaches all the limitations of claim 11, Hastie further teaches:
wherein the first approximate function is a first trained model generated by performing machine learning (Page 343, “where rim = yi − fm−1(xi) is simply the residual of the current model on the ith observation.” Page 356, “The boosted tree model is a sum of such trees, (equation 10.28) – EN: this denotes calling the accumulated approximation (fm-1) the “current model”, which is made up of a sum of previously trained decision trees.) based on the training input data and the training correct answer data. (Page 338, “each of the training observations (xi , yi), i = 1, 2, . . . , N”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the first approximate function being a trained machine learning model of Hastie. The motivation for doing so would be to ensure the foundational baseline model is accurate and capable of capturing complex patterns in the historical training data, rather than relying on a simplistic formula. Hastie describes building the baseline model out of decision trees Page 356, “The boosted tree model is a sum of such trees…” Hastie further refers to this accumulated sum of trained trees as the established baseline model from which errors are calculated “..where rim = yi − fm−1(xi) is simply the residual of the current model on the ith observation” (page 343).
Claim 14
Jacobs in view of Hsieh in view of XGboost further in view of Hastie teaches all the limitations of claim 11, Hastie further teaches:
wherein the first approximate function is a function obtained by formulating a relationship between the training input data and the training correct answer data. (Page 361 - Algorithm 10.3 Gradient Tree Boosting Algorithm, “1. Initialize f0(x) = arg miny PN i=1 L(yi , y)” Page 360, “The first line of the algorithm initializes to the optimal constant model, which is just a single terminal node tree.” – EN: this denotes in the very first step of gradient boosting, the base approximation (f0) is not a fully trained sequential machine learning model yet; it is mathematically derived by minimizing the loss over the data (a mathematical formula).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the first approximate function being a mathematical function between input and answer data of Hastie. The motivation for doing so would be to establish an efficient and mathematically optimal baseline for the system. Hastie describes deriving this initial function not through complex sequential machine learning but through a direct mathematical formula to find the optimal base line, “The first line of the algorithm initializes to the optimal constant model, which is just a single terminal node tree.” (Hastie, page 360).
Claims 6 is rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Jacobs et al. (“Adaptive Mixtures of Local Experts”, hereinafter “Jacobs”) in view of patent application US 2018/0137427 A1 Hsieh et al., hereinafter “Hsieh” in view of non-patent literature XGboost (“xgboost Release 1.0.2”, hereinafter “XGboost”) further in view of Zhou (“Ensemble Methods Foundations and Algorithms”, hereinafter “Zhou”)
Claim 6
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Jacobs further teaches:
wherein the update amount is a value calculated by … multiplying a difference between the inferred output data and the correct answer data by a learning rate by … (Page 80 equation 1.1 and Page 85, “All simulations were performed using a simple gradient descent algorithm with fixed step size t” – EN: this denotes using a fixed step size in a gradient descent algorithm which is synonymous with a learning rate. When a gradient descent algorithm is applied to the error function in equation 1.1, the mathematical derivative extracts the difference term. The algorithm then multiplies this gradient by the step size to calculate the final weight.)
Jacobs in view of Hsieh in view of XGboost does not explicitly disclose:
…dividing a value obtained by … the number of inference models that constitute the ensemble learning-type inference model.
However, Zhou teaches:
…dividing a value obtained by … the number of inference models that constitute the ensemble learning-type inference model. (Zhou Page 68, “Suppose we are given a set of T individual learners {hi,…, hr} … Specifically, simple averaging gives the combined output H(x) as (equation 4.1)”
PNG
media_image3.png
76
885
media_image3.png
Greyscale
Page 76, “If all the individual classifiers are treated equally, the simple soft voting method generates the combined output by simply averaging all the individual outputs, and the final output for class c; is given by (equation 4.23)” – EN: this denotes the concept of equally distributing weight among models using “simple averaging” by multiplying by (1/T) which is equivalent to dividing by T – number of models. )
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the simple averaging technique of Zhou. The motivation for doing so would be to reduce the variance of the predictions, thereby preventing the ensemble model from overfitting to the training data. See Page 71 of Zhou, “In particular, with a large ensemble, there are a lot of weights to learn, and this can easily lead to overfitting; simple averaging does not have to learn any weights, and so suffers little from overfitting. In general, it is widely accepted that simple averaging is appropriate for combining learners with similar performances, whereas if the individual learners exhibit nonidentical strength, weighted averaging with unequal weights may achieve a better performance.”
Claims 15 is rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Jacobs et al. (“Adaptive Mixtures of Local Experts”, hereinafter “Jacobs”) in view of patent application US 2018/0137427 A1 Hsieh et al., hereinafter “Hsieh” in view of non-patent literature XGboost (“xgboost Release 1.0.2”, hereinafter “XGboost”) further in view of Japanese patent application JP2019057016A Nishiyama et al., hereinafter “Nishiyama”.
Claim 15
Jacobs in view of Hsieh in view of XGboost teaches all the limitations of claim 1, Nishiyama teaches:
further comprising a conversion processor configured to convert, when the correct answer data is a label, the label into a numerical value. (Para 37, “Furthermore, the conversion unit 15c converts labels indicating whether the learning data is benign or malignant into numerical labels. For example, the label is expressed as a number, with 0 representing a benign label and 1 representing a malignant label. FIG. 7 illustrates an example of feature vectors converted from feature quantities and numerical labels converted from labels.”)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the conversion unit that converts labels into numerical value of Nishiyama. The motivation for doing so would be to allow the machine learning model to calculate a continuous numerical score for categorical data. Nishiyama states that by converting text labels into numbers, the system can calculate a numerical degree or score for the labels, Para 47 (Nishiyama), “creates a model indicating the degree of benignity or malignancy of the communication log as the classifier” Para 43 (Nishiyama), “determines that the communication log is benign or malignant if the score indicating the degree of benignity or malignancy of the communication log output by the classifier 14a is higher than a predetermined threshold.”
Claims 20 is rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature Jacobs et al. (“Adaptive Mixtures of Local Experts”, hereinafter “Jacobs”) in view of patent application US 2018/0137427 A1 Hsieh et al., hereinafter “Hsieh” in view of non-patent literature XGboost (“xgboost Release 1.0.2”, hereinafter “XGboost”) further in view of patent application US 2022/0056953 A1 Fujimoto et al., hereinafter “Fujimoto”.
Claim 20
Fujimoto teaches:
A control apparatus for controlling a target apparatus, the control apparatus… (Para 89, “In the present embodiment, the control unit 40 controls the position of the shaft 115 by using a machine learning technique.” Para 81, “the control unit 40 is constituted by a microcomputer and a memory device or the like that stores software for causing the microcomputer to operate.”) …acquire input data and correct answer data that corresponds to the input data from the target apparatus; (Para 94, “the control-target device 50 is at least one of the shaft 115 , the magnetic bearings 21 and 22 , and the displacement sensors 31 and 32” Para 96, “The state variable acquisition unit 43 observes the state of the magnetic bearing device 10 while the magnetic bearing device 10 is in operation and acquires information regarding the observed state as a state variable… the state variable is the output values of the displacement sensors 31 and 32” Para 99, “The evaluation data is used as training data in supervised learning.” Para 119, “The learning data is a set of pairs of input data and training data corresponding to the input data.” – EN: this denotes acquiring state variables (input data) from the sensors of the control-target device (target apparatus) and pairing them with corresponding training data (correct answer data) to form a learning dataset.)
The remaining limitations of claim 20 are substantially the same as claim 1, therefore claim 20 is rejected under the same rationale as claim 1.
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the ensemble learning network of Jacobs, the hardware and additional learning of Hsieh, and the trained decision trees of XGboost with the control apparatus for controlling a target apparatus and acquiring data from said target apparatus of Fujimoto. The motivation for doing so would be to dynamically adapt the inference models to real world physical changes such as mechanical wear or environmental shifts in order to maintain long-term operational stability. Fujimoto states that without adapting to the target apparatus data, “appropriate voltage command values are not obtained because of a device-to-device quality variation, a temporal change of the system, and the like. Consequently, the stability of the control… may decrease or the shaft may touch a touchdown bearing” whereas applying this continuous learning allows the system to “maintain the stability of control of the position of the shaft 115 for a long period” (Paras. 187-188).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123