Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Continued Examination Under 37 CFR 1.114
2. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 9/26/2025 has been entered. Claims 1, 2, 8, 9, 15, and 16 have been amended. Claims 1-20 remain pending in the application.
Response to Arguments
3. Applicant’s arguments with respect to claims have been considered but are moot in view of new ground of rejection. See rejections below for details.
Claim Rejections – 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. Claims 1-6, 8-13, and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Moore et al. (U.S. Patent Application Pub. No. US 20200057958 A1) in view of Koch et al. (U.S. Patent Application Pub. No. US 20180240041 A1).
Claim 1: Moore teaches an apparatus, the apparatus comprising:
a processor (i.e. processor; para. [0033]); and
memory comprising instructions that when executed by the processor cause the processor to (i.e. the processor may be coupled to memory, such as RAM, ROM, flash memory, a hard disk or any other device capable of storing electronic information. The memory may store instructions adapted to be executed by the processor to perform the techniques according to embodiments of the disclosed subject matter; para. [0041]):
identify a list of hyperparameters associated with a machine learning (ML) model, the list of hyperparameters comprising hyperparameters ordered based on influence on accuracy of the ML model (i.e. n stage 210, the method 200 may identify and rank the hyperparameters associated with the machine learning model selected in stage 205 according to their respective influence on the one or more performance metrics. This may be achieved using a secondary machine learning model that receives the previously-discussed data resulting from training the selected machine learning model across the plurality of datasets and one or more selected performance metrics and returns a ranking of the associated hyperparameters according to their respective influence on the one or more selected performance metrics. The secondary machine learning model may utilize a random forest algorithm or other conventional machine learning algorithms capable of computing hyperparameter importance in a model; para. [0024]), the list of hyperparameters including a first hyperparameter and a second hyperparameter having less influence on accuracy of the ML model than the first hyperparameter (i.e. a disclosed method may select a machine learning model and train one or more models using one or more datasets. From the one or more subsequently trained models, one or more hyperparameters having a greater influence on the model behavior across one or more of the datasets may be identified and compiled into a list. For each hyperparameter on the list, a range of values may be searched to identify suitable values that cause the machine learning model to perform in accordance with the performance metrics that may be specified by the data scientist. The selected machine learning model may then be configured to train using the identified, suitable hyperparameter values; para. [0019]);
generate a first copy of the ML model for optimization of the first hyperparameter, wherein the first copy corresponds to a first hyperparameter combination and utilizes a value for the second hyperparameter (i.e. fig. 2B, In stage 265, the method 250 may retrieve hyperparameters and their associated values previously-used with the selected machine learning model from the hyperparameter store previously described. The retrieved hyperparameters and their associated values may have been previously used with the same version or a different version as the selected machine learning model. As previously discussed with respect to stage 220, the version of the machine learning algorithm may be stored in the hyperparameter store. The hyperparameter store may also associate the schema of the dataset for which the model was trained. Because the dataset may affect the suitability of the hyperparameters, stage 265 may also compare the schema of the dataset selected in 255 with the schema of the dataset stored in the hyperparameter store to assess the similarities and differences; para. [0003, 0028-0030]);
generate a second copy of the ML model for optimization of the first and second hyperparameters, wherein the second copy corresponds to a second hyperparameter combination (i.e. the hyperparameter values for each previously-used hyperparameter determined to be in common with the hyperparameters of the version of the selected machine learning model may be searched based on a threshold value. For example, if a previously-used hyperparameter value is 10, stage 270 may select a threshold range of 5, so that values between 5 and 15 will be tested for suitability. As previously discussed, the search for suitable hyperparameter values may be carried out using any conventional machine learning algorithm. Preferably, a grid search or equivalent algorithm may be used that allows for specifying the size and/or granularity of the search for each hyperparameter; para. [0003, 0025, 0029]);
optimize the first and second copies of the ML model with an algorithm to produce a first hyperparameter value associated with the first copy of the ML model and first and second hyperparameter values associated with the second copy of the ML model (i.e. Having identified and ranked the hyperparameters according to influence in stage 210, stage 215 may search for suitable hyperparameter values using any conventional machine learning algorithm. Preferably, a grid search algorithm may be used that allows for specifying the size and/or granularity of the search for each hyperparameter. Hyperparameters determined to have a stronger influence on the performance metrics may be searched for suitable values with greater granularity. Hyperparameters determined to have a weaker influence on the performance metrics may be searched for suitable values with lesser granularity. In this way, computing resources may be more efficiently utilized by allocating time for search where the result may be more productive; para. [0025]), wherein the algorithm is configured to generate accuracies (i.e. The machine learning model selected in 205 may have previously trained across a plurality of datasets and may have generated data useful in determining a degree of influence with respect to performance metrics for each hyperparameter associated with the selected machine learning model. The performance metrics may be determined automatically or by a data scientist and may include, for example, accuracy, error, precision, recall, area under the precision-recall curve (AuPR), area under the receiver operating characteristic curve (AuROC), and the like. It should be appreciated that the selection of one or more performance metrics may be relevant in assessing whether one hyperparameter value is better than another, and that no one hyperparameter value may perform better than all other hyperparameter values in view of every performance metric; para. [0023]) for the first and second copies of the ML model (i.e. Having identified and ranked the hyperparameters according to influence in stage 210, stage 215 may search for suitable hyperparameter values using any conventional machine learning algorithm. Preferably, a grid search algorithm may be used that allows for specifying the size and/or granularity of the search for each hyperparameter. Hyperparameters determined to have a stronger influence on the performance metrics may be searched for suitable values with greater granularity. Hyperparameters determined to have a weaker influence on the performance metrics may be searched for suitable values with lesser granularity. In this way, computing resources may be more efficiently utilized by allocating time for search where the result may be more productive; para. [0025]), and the hyperparameter values associated with the first and second copies of the ML model are produced based on the accuracies for the first and second copies of the ML model (i.e. FIG. 2B illustrates an example flow diagram 250 for selecting one or more suitable values for machine learning model hyperparameters. In 255, the method receives a selection of a machine learning model and one or more datasets. The machine learning model may be selected according to method 100 via the secondary machine learning model in stage 115, selected by a user, or selected according to other conventional methods known in the art. The machine learning model selected in stage 255 may have an associated version that may be identified in 260. The version may correspond to the version of the machine learning algorithm that the model employs, for example. A newer version of a machine learning algorithm may utilize new hyperparameters that a prior version lacked and/or may have eliminated other hyperparameters. In general, across multiple versions of a machine learning algorithm, all or a majority of hyperparameters may remain the same, so as to warrant an advantage by storing and recalling previously-used, suitable hyperparameters within the hyperparameter store; para. [0027]);
identify accuracy of the first copy of the ML model using the first hyperparameter value associated with the first copy of the ML model for the first hyperparameter and the default value for the second hyperparameter (i.e. FIG. 2A illustrates an example flow diagram 200 for selecting one or more suitable values for machine learning model hyperparameters. In 205, the method receives a selection of a machine learning model and one or more datasets. The machine learning model may be selected according to method 100 via the secondary machine learning model in stage 115, selected by a user, or selected according to other conventional methods known in the art. The machine learning model selected in 205 may have previously trained across a plurality of datasets and may have generated data useful in determining a degree of influence with respect to performance metrics for each hyperparameter associated with the selected machine learning model. The performance metrics may be determined automatically or by a data scientist and may include, for example, accuracy, error, precision, recall, area under the precision-recall curve (AuPR), area under the receiver operating characteristic curve (AuROC), and the like; para. [0023, 0024]);
identify accuracy of the second copy of the ML model using the first hyperparameter value associated with the second copy of the ML model for the first hyperparameter and the second hyperparameter value associated with the second copy of the ML model for the second hyperparameter (i.e. FIG. 2A illustrates an example flow diagram 200 for selecting one or more suitable values for machine learning model hyperparameters. In 205, the method receives a selection of a machine learning model and one or more datasets. The machine learning model may be selected according to method 100 via the secondary machine learning model in stage 115, selected by a user, or selected according to other conventional methods known in the art. The machine learning model selected in 205 may have previously trained across a plurality of datasets and may have generated data useful in determining a degree of influence with respect to performance metrics for each hyperparameter associated with the selected machine learning model. The performance metrics may be determined automatically or by a data scientist and may include, for example, accuracy, error, precision, recall, area under the precision-recall curve (AuPR), area under the receiver operating characteristic curve (AuROC), and the like; para. [0023, 0024]);
determine the first hyperparameter combination results in a more accurate ML model than the second hyperparameter combination (i.e. Having identified and ranked the hyperparameters according to influence in stage 210, stage 215 may search for suitable hyperparameter values using any conventional machine learning algorithm. Preferably, a grid search algorithm may be used that allows for specifying the size and/or granularity of the search for each hyperparameter. Hyperparameters determined to have a stronger influence on the performance metrics may be searched for suitable values with greater granularity. Hyperparameters determined to have a weaker influence on the performance metrics may be searched for suitable values with lesser granularity; para. [0025]); and
create a production ML model based on the first hyperparameter combination (i.e. In stage 280, the machine learning model selected in 255 may be trained using the dataset selected in 255 and the selected suitable hyperparameter values determined in stage 275; para. [0030]).
Moore does not explicitly teach utilizes a default value for the second hyperparameter; wherein the second hyperparameter is nonoptimizable when optimizing the first copy of the ML model; optimize the ML model with a genetic algorithm, wherein the genetic algorithm is configured to generate accuracies for the ML model, and the hyperparameter values associated with the ML model are produced based on the accuracies for the ML model.
However, Koch teaches generate a first copy of the ML model for optimization of the first hyperparameter (i.e. Each session of the plurality of sessions executes training and scoring of the model type using the input dataset in parallel with other sessions of the plurality of sessions … For each session of the plurality of sessions, a hyperparameter configuration is assigned to the session of the plurality of sessions; para. [0005, 0041, 0053]), wherein the first copy corresponds to a first hyperparameter combination and utilizes a default value for the second hyperparameter (i.e. A hyperparameter configuration includes a value for each hyperparameter of the plurality of hyperparameters … the user may identify one or more of the hyperparameters to exclude from the evaluation such that a single value is used for that hyperparameter when selecting values for each hyperparameter configuration. When a hyperparameter is excluded, a default value defined for the hyperparameter may be used for each hyperparameter configuration; para. [0005, 0089, 0097]), wherein the second hyperparameter is nonoptimizable when optimizing the first copy of the ML model (i.e. the user may identify one or more of the hyperparameters to exclude from the evaluation … EXCLUDE indicates whether or not to exclude the hyperparameter from the tuning evaluation; para. [0089, 0097]);
generate a second copy of the ML model for optimization of the first and second hyperparameters (i.e. A plurality of hyperparameter configurations is determined using a search method of the search method type … Each hyperparameter configuration of the plurality of hyperparameter configurations is unique; para. [0005, 0053]), wherein the second copy corresponds to a second hyperparameter combination (i.e. A plurality of hyperparameter configurations is determined using a search method of the search method type … Each hyperparameter configuration of the plurality of hyperparameter configurations is unique; para. [0005, 0049]);
optimize the first and second copies of the ML model with a genetic algorithm (i.e. the GA search method defines a family of local search algorithms that seek optimal solutions to problems by applying the principles of natural selection and evolution. A GA search method can be applied to almost any optimization problem; para. [0136, 0137, 0182]) to produce a first hyperparameter value associated with the first copy of the ML model (i.e. the hyperparameter configuration defines a value for each hyperparameter for the model type; para. [0005, 0049]) and first and second hyperparameter values associated with the second copy of the ML model (i.e. A value for each of these hyperparameters is defined in each hyperparameter configuration for the neural network model type; para. [0005, 0094]), wherein the genetic algorithm is configured to generate accuracies for the first and second copies of the ML model (i.e. model scoring uses first validation dataset subset 416 and/or Pth validation dataset subset 436 to determine how well the generated model performed … The objective function specifies a measure of model error (performance) to be used to identify a best configuration of the hyperparameters among those evaluated; para. [0041, 0100]), and the hyperparameter values associated with the first and second copies of the ML model are produced based on the accuracies for the first and second copies of the ML model (i.e. Based on the results and the current tuning search method(s), iteration manager 314 determines a next set of hyperparameter configurations to evaluate in a next iteration. The best model hyperparameter configurations from the previous iteration are used to generate the next population of hyperparameter configurations to evaluate with the selected mode type; para. [0181]);
identify accuracy of the first copy of the ML model using the first hyperparameter value associated with the first copy of the ML model for the first hyperparameter and the default value for the second hyperparameter (i.e. model scoring uses first validation dataset subset 416 and/or Pth validation dataset subset 436 to determine how well the generated model performed … The objective function specifies a measure of model error (performance) to be used to identify a best configuration of the hyperparameters among those evaluated; para. [0041, 0089, 0100]);
identify accuracy of the second copy of the ML model using the first hyperparameter value associated with the second copy of the ML model for the first hyperparameter and the second hyperparameter value associated with the second copy of the ML model for the second hyperparameter (i.e. The model is trained using the assigned hyperparameter configuration and a training dataset that is a first portion of the input dataset. The trained model is scored using the assigned hyperparameter configuration and a validation dataset that is a second portion of the input dataset; para. [0005, 0094])
determine the first hyperparameter combination results in a more accurate ML model than the second hyperparameter combination (i.e. A best hyperparameter configuration is identified based on an extreme value of the stored objective function values. The identified best hyperparameter configuration is output; para. [0005, 0073, 0188]); and
create a production ML model based on the first hyperparameter combination (i.e. a trained model output that includes information to execute the model generated using the input dataset with the best hyperparameter configuration. For example, the trained model output includes information to execute the model generated using the input dataset with the best hyperparameter configuration that may be saved in selected model data 320 and used to score a second dataset; para. [0073, 0188, 0191]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 2: Moore and Koch teach the apparatus of claim 1. Moore further teaches wherein the instructions, when executed by the processor, further cause the processor to generate a list of hyperparameter combinations comprising the first and second hyperparameter combinations (i.e. Having identified and ranked the hyperparameters according to influence in stage 210, stage 215 may search for suitable hyperparameter values using any conventional machine learning algorithm. Preferably, a grid search algorithm may be used that allows for specifying the size and/or granularity of the search for each hyperparameter. Hyperparameters determined to have a stronger influence on the performance metrics may be searched for suitable values with greater granularity. Hyperparameters determined to have a weaker influence on the performance metrics may be searched for suitable values with lesser granularity. In this way, computing resources may be more efficiently utilized by allocating time for search where the result may be more productive; para. [0025, 0029]) with the hyperparameter combinations ordered based on accuracy of the first and second copies of the ML model (i.e. In stage 210, the method 200 may identify and rank the hyperparameters associated with the machine learning model selected in stage 205 according to their respective influence on the one or more performance metrics. This may be achieved using a secondary machine learning model that receives the previously-discussed data resulting from training the selected machine learning model across the plurality of datasets and one or more selected performance metrics and returns a ranking of the associated hyperparameters according to their respective influence on the one or more selected performance metrics. The secondary machine learning model may utilize a random forest algorithm or other conventional machine learning algorithms capable of computing hyperparameter importance in a model; para. [0024]), wherein the first copy of the ML model is more accurate than the second copy of the ML model (i.e. a greater number of hyperparameter values may be tested for suitability where it may be less certain that those previously-used hyperparameter values will be suitable for the present use. In stage 280, the machine learning model selected in 255 may be trained using the dataset selected in 255 and the selected suitable hyperparameter values determined in stage 275; para. [0030]).
Moore does not explicitly teach wherein the first hyperparameter combination is a top ranked hyperparameter combination in the list of hyperparameter combinations.
However, Koch further teaches wherein the instructions, when executed by the processor, further cause the processor to generate a list of hyperparameter combinations comprising the first and second hyperparameter combinations (i.e. A plurality of hyperparameter configurations is determined using a search method of the search method type. A hyperparameter configuration includes a value for each hyperparameter of the plurality of hyperparameters. Each hyperparameter configuration of the plurality of hyperparameter configurations is unique; para. [0005, 0073]) with the hyperparameter combinations ordered based on accuracy of the first and second copies of the ML model (i.e. a “Tuner Results” output table that includes a default configuration and up to ten of the best hyperparameter configurations (based on an extreme (minimum or maximum) objective function value) identified, where each configuration listed includes the hyperparameter values and objective function value for comparison; para. [0073]), wherein the first copy of the ML model is more accurate than the second copy of the ML model (i.e. A best hyperparameter configuration is identified based on an extreme value of the stored objective function values. The identified best hyperparameter configuration is output; para. [0005]); and wherein the first hyperparameter combination is a top ranked hyperparameter combination in the list of hyperparameter combinations (i.e. a “Tuner Results” output table that includes a default configuration and up to ten of the best hyperparameter configurations (based on an extreme (minimum or maximum) objective function value) identified, where each configuration listed includes the hyperparameter values and objective function value for comparison; a “Tuner Evaluation History” output table that includes all of the hyperparameter configurations evaluated, where each configuration listed includes the hyperparameter values and objective function value for comparison; a “Best Configuration” output table that includes values of the hyperparameters and the objective function value for the best configuration identified; para. [0073, 0188]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 3: Moore and Koch teach the apparatus of claim 2. Moore further teaches wherein the list of hyperparameters includes a third hyperparameter ordered between the first and second hyperparameters and the list of hyperparameter combinations includes a third hyperparameter combinations ordered below the first and second hyperparameter combinations (i.e. In stage 210, the method 200 may identify and rank the hyperparameters associated with the machine learning model selected in stage 205 according to their respective influence on the one or more performance metrics. This may be achieved using a secondary machine learning model that receives the previously-discussed data resulting from training the selected machine learning model across the plurality of datasets and one or more selected performance metrics and returns a ranking of the associated hyperparameters according to their respective influence on the one or more selected performance metrics. The secondary machine learning model may utilize a random forest algorithm or other conventional machine learning algorithms capable of computing hyperparameter importance in a model; para. [0024]).
However, Koch further teaches the list of hyperparameter combinations includes a third hyperparameter combinations ordered below the first and second hyperparameter combinations (i.e. a “Tuner Results” output table that includes a default configuration and up to ten of the best hyperparameter configurations (based on an extreme (minimum or maximum) objective function value) identified, where each configuration listed includes the hyperparameter values and objective function value for comparison; a “Tuner Evaluation History” output table that includes all of the hyperparameter configurations evaluated, where each configuration listed includes the hyperparameter values and objective function value for comparison; a “Best Configuration” output table that includes values of the hyperparameters and the objective function value for the best configuration identified; para. [0073, 0188]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 4: Moore and Koch teach the apparatus of claim 1. Moore does not explicitly teach classify data with the production ML model.
However, Koch further teaches classify data with the production ML model (i.e. Prediction application 1822 may be implemented as a Web application. Prediction application 1822 may be integrated with other system processing tools to automatically process data generated as part of operation of an enterprise, to classify data in the processed data, to identify any outliers in the processed data, and/or to provide a warning or alert associated with the data classification and/or outlier identification; para. [0245]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 5: Moore and Koch teach the apparatus of claim 1. Moore does not explicitly teach to simultaneously optimize the first and second copies of the ML model with the algorithm.
However, Koch further teaches to simultaneously optimize the first and second copies of the ML model with the algorithm (i.e. Each session of the plurality of sessions executes training and scoring of the model type using the input dataset in parallel with other sessions of the plurality of sessions. A plurality of hyperparameter configurations is determined using a search method of the search method type. A hyperparameter configuration includes a value for each hyperparameter of the plurality of hyperparameters. Each hyperparameter configuration of the plurality of hyperparameter configurations is unique; para. [0005]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 6: Moore and Koch teach the apparatus of claim 1. Moore further teaches wherein the production ML model utilizes the first hyperparameter value associated with the first copy of the ML model for the first hyperparameter and the default value for the second hyperparameter (i.e. fig. 2B, In stage 265, the method 250 may retrieve hyperparameters and their associated values previously-used with the selected machine learning model from the hyperparameter store previously described. The retrieved hyperparameters and their associated values may have been previously used with the same version or a different version as the selected machine learning model. As previously discussed with respect to stage 220, the version of the machine learning algorithm may be stored in the hyperparameter store. The hyperparameter store may also associate the schema of the dataset for which the model was trained. Because the dataset may affect the suitability of the hyperparameters, stage 265 may also compare the schema of the dataset selected in 255 with the schema of the dataset stored in the hyperparameter store to assess the similarities and differences; para. [0003, 0028-0030]).
However, Koch further teaches wherein the production ML model utilizes the first hyperparameter value associated with the first copy of the ML model for the first hyperparameter and the default value for the second hyperparameter (i.e. A hyperparameter configuration includes a value for each hyperparameter of the plurality of hyperparameters … the user may identify one or more of the hyperparameters to exclude from the evaluation such that a single value is used for that hyperparameter when selecting values for each hyperparameter configuration. When a hyperparameter is excluded, a default value defined for the hyperparameter may be used for each hyperparameter configuration; para. [0005, 0089, 0097]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Moore to include the feature of Koch. One would have been motivated to make this modification because it reduces search space and computation while objectively identifying the best hyperparameter configuration for training the final model.
Claim 8 is similar in scope to Claim 1 and is rejected under a similar rationale.
Moore teaches at least one non-transitory computer-readable medium comprising a set of instructions that, in response to being executed by a processor circuit (i.e. processor; para. [0033]), cause the processor circuit to (i.e. the processor may be coupled to memory, such as RAM, ROM, flash memory, a hard disk or any other device capable of storing electronic information. The memory may store instructions adapted to be executed by the processor to perform the techniques according to embodiments of the disclosed subject matter; para. [0041]).
Claims 9-13 are similar in scope to Claims 2-6 and are rejected under a similar rationale.
Claims 15-19 are similar in scope to Claims 1, 2, 4-6 and are rejected under a similar rationale.
6. Claims 7, 14, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Moore in view of Koch, and further in view of Arjasakusuma et al. (Evaluating Variable Selection and Machine Learning Algorithms for Estimating Forest Heights by Combining Lidar and Hyperspectral Data, ISPRS Int. J. Geo-Inf, published 2020, pages 1-26).
Claim 7: Moore and Kawajiri teach the apparatus of claim 1. Moore further teaches wherein the instructions, when executed by the processor, further cause the processor to generate the list of hyperparameters associated with the ML model with a feature algorithm (i.e. In stage 210, the method 200 may identify and rank the hyperparameters associated with the machine learning model selected in stage 205 according to their respective influence on the one or more performance metrics. This may be achieved using a secondary machine learning model that receives the previously-discussed data resulting from training the selected machine learning model across the plurality of datasets and one or more selected performance metrics and returns a ranking of the associated hyperparameters according to their respective influence on the one or more selected performance metrics. The secondary machine learning model may utilize a random forest algorithm or other conventional machine learning algorithms capable of computing hyperparameter importance in a model; para. [0024]).
Moore does not explicitly teach feature importance algorithm.
However, Arjasakusuma teaches feature importance algorithm (i.e. Feature selection analyses and dimensionality reduction were conducted using Boruta (BO), Simulated Annealing (SA), Genetic Algorithm (GA), and Principal Component Analysis (PCA). These methods were used to select the significant variables from AISA hyperspectral bands (479 variables), lidar statistical metrics (37 variables), and the combination of both datasets (516 variables) for modeling the forest height data; page 10).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Moore and Koch to include the feature of Arjasakusuma. One would have been motivated to make this modification because the system can more accurately determine which hyperparameters have the most significant impact on the model’s performance.
Claims 14 and 20 are similar in scope to Claim 7 and are rejected under a similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Loh et al. (Pub. No. US 20200057944 A1), the optimization apparatus 100 may exclude, for each cluster, hyperparameter samples whose evaluation scores are less than a threshold and construct the optimal hyperparameter sample set based on the remaining hyperparameter samples. According to an embodiment, the process of excluding some hyperparameter samples based on evaluation scores may be performed before the clustering operation S140.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TAN H TRAN/Primary Examiner, Art Unit 2141