Prosecution Insights
Last updated: October 02, 2026
Application No. 18/501,886

METHODS AND SYSTEMS FOR DEVELOPING DECISION TREE MACHINE LEARNING MODELS

Non-Final OA §101§103
Filed
Nov 03, 2023
Priority
May 17, 2023 — IN 202321034599
Examiner
ADMASU, MAHLIET TASEW
Art Unit
Tech Center
Assignee
Capital One Services LLC
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
6m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
15 currently pending
Career history
12
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
62.0%
+22.0% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
CTNF 18/501,886 CTNF 101331 DETAILED ACTION This communication is in response to the Application No. 18/501,886 filed November 03, 2023 in which Claims 1 - 20 are presented for examination. Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Specification 07-29 AIA The disclosure is objected to because of the following informalities: Par. [0074] ends with the phrase “In some embodiments, a missing-informative” which is an unfinished sentence . Appropriate correction is required. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-20 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a computer-implemented method type claim. Therefore, Claims 1-7 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. determining a set of numeric variables for each of the plurality of categorical variables (mental process – determining a set of numeric variables for each of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to determine a set of numeric variables for each of the plurality of categorical variables based on said analysis) determining an event rate for each of the plurality of categorical variables (mathematical concept – according to the specification, Par. [0040], “In general, an event rate is a measure of how often a statistical event occurs within a data set or experimental group. The category-wise event rate process needs to be calculated recursively for each node since the population composition changes from node to node”, therefore determining the event rate involves statistical calculation and evaluation of event-occurrence frequencies within categorical variable groupings, which is a mathematical relationship and statistical analysis) determining, using training data, a plurality of splits for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, wherein the plurality of splits comprises a plurality of multi-categorical splits assigning multiple of the plurality of categorical variables to a single node based on the event rate ( mental process – determining a plurality of splits for assigning the plurality of categorical variables using training data may be performed mentally or using pen and paper by a user observing or analyzing the training data and using judgment/evaluation to determine the plurality of splits and assignment of categorical variables based on the event rate) […] to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes assigned one of the plurality of multi-categorical splits (mental process – generating a machine learning decision tree model may be performed mentally or using pen and paper by a user using judgment/evaluation to generate a simple machine learning/decision tree model based on the analyzed data and selected splits) […] to generate a refined trained model (mental process – generating a refined trained model may be performed by a user using pen and paper to generate a simple refined trained model) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] at least one processor of a computing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] ( Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. […] at least one processor of a computing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2 - 7. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. determining the multi-categorical splits based on a grouping threshold of the event rate associated with each of the plurality of categorical variables (mental process – determining the multi-categorical splits based on a grouping threshold of the event rate may be performed mentally by a user observing/analyzing the grouping threshold of the event rate associated with each of the plurality of categorical variables and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on. determining a specified depth of the machine learning decision tree model (mental process – determining a specified depth of the machine learning decision tree model may be performed mentally or using pen and paper by a user observing/analyzing the decision tree structure and using judgment/evaluation to determine how many levels or nodes the decision tree model should contain) and determining the multi-categorical splits based on the specified depth (mental process - determining the multi-categorical splits may be performed mentally by a user observing/analyzing the specified depth and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. replacing each of the plurality of categorical variables with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the plurality of categorical variables (mental process – replacing each of the plurality of categorical variables with a set of a plurality of dummy variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to replace each of the plurality of categorical variables with a set of a plurality of dummy variables based on said analysis ) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on. separately determining the plurality of splits for the plurality of categorical variables for each of the plurality of dummy variables (mental process – determining the plurality of splits for the plurality of categorical variables for each of the plurality of dummy variables may be performed mentally by a user observing/analyzing the plurality of categorical variables for each of the plurality of dummy variables and accordingly using judgment/evaluation to determine the plurality of splits separately based on said analysis ) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on. determining missing values of the plurality of categorical variables (mental process – determining missing values of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgement/evaluation to determine missing values based on said analysis) and determining whether the missing values are informative based on a statistical significance of a category of the plurality of categorical variables (mental process - determining whether the missing values are informative based on a statistical significance a category of the plurality of categorical variables may be performed mentally by a user observing/analyzing the statistical significance of a category of the plurality of categorical variables and accordingly using judgment/evaluation based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 6 above, which Claim 7 depends on. imputing the missing values responsive to the missing values being informative (mental process – imputing missing values responsive to the missing values being informative may be performed mentally or using pen and paper by a user observing/analyzing available data and using judgment/evaluation to estimate replacement values for the missing data) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 8: Step 1: Claim 8 is an apparatus type claim. Therefore, Claims 8 - 14 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. determine a set of numeric variables for each of the plurality of categorical variables (mental process – determining a set of numeric variables for each of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to determine a set of numeric variables for each of the plurality of categorical variables based on said analysis) determine an event rate for each of the plurality of categorical variables (mathematical concept – according to the specification, Par. [0040], “In general, an event rate is a measure of how often a statistical event occurs within a data set or experimental group. The category-wise event rate process needs to be calculated recursively for each node since the population composition changes from node to node”, therefore determining the event rate involves statistical calculation and evaluation of event-occurrence frequencies within categorical variable groupings, which is a mathematical relationship and statistical analysis) determine, using training data, a plurality of splits for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, wherein the plurality of splits comprises a plurality of multi-categorical splits assigning multiple of the plurality of categorical variables to a single node based on the event rate ( mental process – determining a plurality of splits for assigning the plurality of categorical variables using training data may be performed mentally or using pen and paper by a user observing or analyzing the training data and using judgment/evaluation to determine the plurality of splits and assignment of categorical variables based on the event rate) […] to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes assigned one of the plurality of multi-categorical splits (mental process – generating a machine learning decision tree model may be performed mentally or using pen and paper by a user using judgment/evaluation to generate a simple machine learning/decision tree model based on the analyzed data and selected splits) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] at least one processor of a computing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] ( Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. at least one processor; and a memory […] (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 8 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 9 - 14. The additional limitations of the dependent claims are addressed below. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 9 depends on. determine the multi-categorical splits based on a grouping threshold of the event rate associated with each of the plurality of categorical variables (mental process – determining the multi-categorical splits based on a grouping threshold of the event rate may be performed mentally by a user observing/analyzing the grouping threshold of the event rate associated with each of the plurality of categorical variables and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 10 depends on. determine a specified depth of the machine learning decision tree model (mental process – determining a specified depth of the machine learning decision tree model may be performed mentally or using pen and paper by a user observing/analyzing the decision tree structure and using judgment/evaluation to determine how many levels or nodes the decision tree model should contain) and determine the multi-categorical splits based on the specified depth (mental process - determining the multi-categorical splits may be performed mentally by a user observing/analyzing the specified depth and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 11 depends on. replace each of the plurality of categorical variables with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the plurality of categorical variables (mental process – replacing each of the plurality of categorical variables with a set of a plurality of dummy variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to replace each of the plurality of categorical variables with a set of a plurality of dummy variables based on said analysis ) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 12: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 12 depends on. separately determine the plurality of splits for the plurality of categorical variables for each of the plurality of dummy variables (mental process – determining the plurality of splits for the plurality of categorical variables for each of the plurality of dummy variables may be performed mentally by a user observing/analyzing the plurality of categorical variables for each of the plurality of dummy variables and accordingly using judgment/evaluation to determine the plurality of splits separately based on said analysis ) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 13: Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 13 depends on. determine missing values of the plurality of categorical variables (mental process – determining missing values of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgement/evaluation to determine missing values based on said analysis) and determine whether the missing values are informative based on a statistical significance of a category of the plurality of categorical variables (mental process - determining whether the missing values are informative based on a statistical significance a category of the plurality of categorical variables may be performed mentally by a user observing/analyzing the statistical significance of a category of the plurality of categorical variables and accordingly using judgment/evaluation based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 14: Step 2A Prong 1: See the rejection of Claim 13 above, which Claim 14 depends on. impute the missing values responsive to the missing values being informative (mental process – imputing missing values responsive to the missing values being informative may be performed mentally or using pen and paper by a user observing/analyzing available data and using judgment/evaluation to estimate replacement values for the missing data) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 15: Step 1: Claim 15 is a non-transitory computer-readable medium type claim. Therefore, Claims 16 - 20 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. determine a set of numeric variables for each of the plurality of categorical variables (mental process – determining a set of numeric variables for each of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to determine a set of numeric variables for each of the plurality of categorical variables based on said analysis) determine an event rate for each of the plurality of categorical variables (mathematical concept – according to the specification, Par. [0040], “In general, an event rate is a measure of how often a statistical event occurs within a data set or experimental group. The category-wise event rate process needs to be calculated recursively for each node since the population composition changes from node to node”, therefore determining the event rate involves statistical calculation and evaluation of event-occurrence frequencies within categorical variable groupings, which is a mathematical relationship and statistical analysis) determine, using training data, a plurality of splits for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, wherein the plurality of splits comprises a plurality of multi-categorical splits assigning multiple of the plurality of categorical variables to a single node based on the event rate ( mental process – determining a plurality of splits for assigning the plurality of categorical variables using training data may be performed mentally or using pen and paper by a user observing or analyzing the training data and using judgment/evaluation to determine the plurality of splits and assignment of categorical variables based on the event rate) […] to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes assigned one of the plurality of multi-categorical splits (mental process – generating a machine learning decision tree model may be performed mentally or using pen and paper by a user using judgment/evaluation to generate a simple machine learning/decision tree model based on the analyzed data and selected splits) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] one or more processors of a computing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] ( Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. […] at least one processor of a computing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) and accessing data […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 15 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 16 - 20. The additional limitations of the dependent claims are addressed below. Regarding Claim 16: Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 16 depends on. determine the multi-categorical splits based on a grouping threshold of the event rate associated with each of the plurality of categorical variables (mental process – determining the multi-categorical splits based on a grouping threshold of the event rate may be performed mentally by a user observing/analyzing the grouping threshold of the event rate associated with each of the plurality of categorical variables and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 17: Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 17 depends on. determine a specified depth of the machine learning decision tree model (mental process – determining a specified depth of the machine learning decision tree model may be performed mentally or using pen and paper by a user observing/analyzing the decision tree structure and using judgment/evaluation to determine how many levels or nodes the decision tree model should contain) and determine the multi-categorical splits based on the specified depth (mental process - determining the multi-categorical splits may be performed mentally by a user observing/analyzing the specified depth and accordingly using judgment/evaluation to determine the multi-categorical splits based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 18: Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 18 depends on. replace each of the plurality of categorical variables with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the plurality of categorical variables (mental process – replacing each of the plurality of categorical variables with a set of a plurality of dummy variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgment/evaluation to replace each of the plurality of categorical variables with a set of a plurality of dummy variables based on said analysis ) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 19: Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 19 depends on. determine missing values of the plurality of categorical variables (mental process – determining missing values of the plurality of categorical variables may be performed mentally by a user observing/analyzing the plurality of categorical variables and accordingly using judgement/evaluation to determine missing values based on said analysis) and determine whether the missing values are informative based on a statistical significance of a category of the plurality of categorical variables (mental process - determining whether the missing values are informative based on a statistical significance a category of the plurality of categorical variables may be performed mentally by a user observing/analyzing the statistical significance of a category of the plurality of categorical variables and accordingly using judgment/evaluation based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 20: Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 20 depends on. impute the missing values responsive to the missing values being informative (mental process – imputing missing values responsive to the missing values being informative may be performed mentally or using pen and paper by a user observing/analyzing available data and using judgment/evaluation to estimate replacement values for the missing data) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 1-2, 8-9, and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Gulin et al. (hereafter Gulin) (US 2019164085) in view of Guillame et al. (hereinafter Guillame) (US 2019251468) . Regarding Claim 1, Gulin teaches a computer-implemented method to generate a machine learning decision tree model (Gulin, Par. [0002], “methods for generating a prediction model”, thus a method is disclosed) for a plurality of categorical variables, comprising, via at least one processor (Gulin, Par. [0193], “a computer or processor”, thus at least one processor is disclosed) of a computing device: determining a set of numeric variables for each of the plurality of categorical variables (Gulin, Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list; and (ii) a number of pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value in the respective ordered list.”, thus determining a set of numeric variables for each of the plurality of categorical variables is disclosed because Gulin teaches generating a numeric representation for a categorical feature using occurrence counts and event-outcome counts associated with that categorical feature. Gulin’s categorical feature corresponds to the plurality of categorical variables, while the generated numeric representation based on occurrences and event outcomes corresponds to the set of numeric variables determined for each categorical variable) determining an event rate for each of the plurality of categorical variables (Gulin, Par. [0065], “The MLA translates the value of the categorical feature 122 (i.e. “Pop”) into a numeric feature using a formula: Count = NumberWINs /NumberOCCURENCEs”, & Par. [0103], “generating the numeric representation thereof comprises: (i) using as the number of total occurrences of the at least one preceding training object with the same categorical feature value: a number of total occurrences of the at least one preceding training object having both the first categorical feature value and the second categorical feature value; and (ii) using as the number of the pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value: a number of the pre-determined outcomes of events associated with at least one preceding training object having both the first categorical feature value and the second categorical feature value”, thus determining an event rate for each of the plurality of categorical variables is disclosed because Gulin teaches calculating a count value as a ratio of NumberWINs to NumberOCCURENCEs for categorical features, where the wins correspond to event outcomes and the occurrences correspond to occurrences of the categorical feature values. The ratio generated from event outcomes and occurrences corresponds to an event rate associated with each categorical variable) determining, using training data, […] for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, […] assigning multiple of the plurality of categorical variables to a single node based on the event rate (Gulin, Par. [0093], “the method comprises: accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0082], “For each node of the decision tree, the MLA translates the categorical features (or groups of categorical features, as the case may be) into numeric representation thereof as has been described above”, & Par. [0076], “For example, rather than just looking at the genre of the song, the formula can analyze co-occurrence of the given genre and the given singer (as examples of two categorical features or a group of categorical features that can be associated with a single training object)”, & Par. [0078], “Where both the Number.sub.WINs (F1 and F2) and the Number.sub.OCCURENCEs(F1 and F2) are considering the wins and co-occurrences of the group of features (F1 and F2) values that occur before the current occurrence of the group of features in the ordered list of training objects 102”, thus determining, using training data, for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, and assigning multiple categorical variables to a single node based on the event rate is disclosed because Gulin teaches accessing training objects containing event indicators and categorical features, building a decision tree, and testing which feature and split value to place at each node. Gulin’s training objects correspond to the training data, the categorical features correspond to the plurality of categorical variables, and the decision-tree nodes correspond to the plurality of nodes. Gulin further teaches using groups of categorical features, such as genre and singer/F1 and F2, and calculating wins and occurrences for the grouped features. The wins correspond to event outcomes, the occurrences correspond to feature value occurrences, and the wins/occurrences relationship corresponds to the event rate used to assign multiple categorical variables to a node) and accessing data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes […] (Gulin, Par. [0093], “accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0221], “The tree model 600 may be referred to as a machine-learning model or a portion of a machine-learning model (e.g., for implementations wherein the machine-learning model relies on multiple tree models). In some instances, the tree model 600 may be referred as a prediction model or a portion of a prediction model (e.g., for implementations wherein the prediction model relies on multiple tree models)”, thus accessing data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes is disclosed because Gulin teaches accessing training objects from a non-transitory computer-readable medium, building a decision tree, and testing which feature and split value to place at the node of each level. Gulin’s training objects correspond to the accessed data, the tree model corresponds to the machine learning decision tree model, and the nodes/levels of the decision tree correspond to the plurality of nodes. The placement of a feature and split value at the node teaches that at least a portion of the nodes are assigned feature/split information during generation of the decision-tree model) Gulin does not explicitly disclose […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] and […] assigned one of the plurality of multi-categorical splits. However, Guillame teaches: […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] (Guillame, Par. [0010], “receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus a plurality of splits wherein the plurality of splits comprises a plurality of multi-categorical splits is disclosed because Guillaume teaches receiving and selecting from a plurality of proposed splits generated from attributes of a training dataset. Guillaume further teaches categorical split conditions for categorical columns and teaches that a super split can refer to a set of splits mapped one-to-one with open leaves of a tree. Guillaume’s plurality of proposed splits corresponds to the plurality of splits, while the super split/set of splits for categorical columns corresponds to the plurality of multi-categorical splits) […] assigned one of the plurality of multi-categorical splits (Guillame, Par. [0010], “The method includes, for each of a plurality of iterations, receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits. The method includes, for each of the plurality of iterations, updating, by the one or more computing devices, a node structure of the decision tree based at least in part on the selected final split”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus assigned one of the plurality of multi-categorical splits is disclosed because Guillaume teaches receiving a plurality of proposed splits, selecting a final split from the plurality of proposed splits, and updating the node structure of the decision tree based on the selected final split. Guillaume further teaches categorical split conditions and teaches that a “super split” can refer to a set of splits mapped one-to-one with open leaves of the tree. Guillaume’s categorical/super splits correspond to the multi-categorical splits, while the updating of the node structure based on the selected final split teaches assigning one of the multi-categorical splits to the decision-tree structure/open leaves) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Gulin’s teaching of event rate based grouped categorical feature processing with Guillaume’s teaching of plurality of proposed splits, super splits, and updating a node structure based on a selected final split. Gulin teaches optimizing which feature and split value to place at each node, while Guillaume teaches an exact distributed split search approach that provides lower volatile memory usage and smaller space, disk, and network complexity. Therefore, a POSITA would have been motivated to incorporate Guillaume’s split generation and node update techniques into Gulin’s decision-tree framework to improve split selection efficiency and reduce computational resource usage during decision tree training (Guillaume, Par. [0023], “the present disclosure provides an exact distributed algorithm to train Random Forest models as well as other decision forest models without relying on approximating best split search”, &Par. [0036], “only a part of the mapping needs to reside in volatile memory at any instant, which advantageously provides lower volatile memory usage”, & Par. [0038], “The methods described herein stand out from existing exact distributed approaches by a smaller space, disk and network complexity”) Regarding Claim 2, Gulin combined with Guillaume teaches all the limitations of claim 1 as cited above and Guillaume further teaches: determining the multi-categorical splits […] associated with each of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, & Par. [0077], “An example listing is given in Algorithm 7. The iteration over the samples can be trivially parallelized (multithreading over sharding). TABLE-US-00002 Algorithm 7 Find the best supersplits for categorical attribute j and tree p. Nodes are open when they are still subject to splitting - typically nodes are closed when they reach some purity level or when their cardinal is below some threshold. H.sub.h ∈ [1,l] is an empty bi-histogram between the labels and the attribute j for the leaf l for all i in 1,...,n // This loop can be parallelized do h ← sample2node(i) if h is a closed node then continue if candidate feature(j,h,p) is false then continue B ← bag(i,p) // Number of times i is sampled in tree p if B = 0 then continue Add (x.sub.i,j,y.sub.i) weighted by B to H.sub.h end for for all open leaf h do Find best condition using bi- histogram H.sub.h end for”, thus determining the multi-categorical splits associated with each of the plurality of categorical variables is disclosed because Guillaume teaches computing bi-variate histograms between categorical attribute values and label values and identifying optimal split conditions for categorical attributes. Guillaume further teaches finding “best supersplits” for categorical attributes using bi-histograms associated with leaves/nodes of the tree. Guillaume’s categorical attributes correspond to the categorical variables, while the identified optimal split conditions and supersplits correspond to the multi-categorical splits associated with the categorical variables) Guillaume does not explicitly teach […] based on a grouping threshold of the event rate […]. However, Gulin teaches […] based on a grouping threshold of the event rate […] (Gulin, Par. [0148], “The method comprises: generating a range of all possible values of the categorical features; applying a grid to the range to separate the range into region, each region having a boundary; using the boundary as the split value; the generating and the applying being executed before the categorical feature value is translated into the numeric representation thereof”, & Par. [0088], “In a specific embodiment of the present technology, the MLA generates the splits by generating a range of all possible values for the splits (for a given counter having been generated based on the given categorical feature) and applying a pre-determined grid. In some embodiments of the present technology, the range can be between 0 and 1. In other embodiments of the present technology, which is especially pertinent when a coefficient (R.sub.constant) is applied to calculating the values of the counts, the range can be between: (i) the value of the coefficient and (ii) the value of coefficient plus one”, thus based on a grouping threshold of the event rate is disclosed because Gulin teaches generating ranges of values for categorical features/splits, applying a grid to separate the range into regions, and using region boundaries as split values. Gulin further teaches that the split ranges may be generated from counts associated with categorical features. The counts generated from categorical feature occurrences and event outcomes correspond to event rate information, while the region boundaries/split values generated from the grid correspond to grouping thresholds used for determining the splits) Regarding Claim 8, Gulin teaches an apparatus comprising: at least one processor (Gulin, Par. [0193], “a computer or processor”, thus at least one processor is disclosed) ; and a memory (Gulin, Par. [0170], “memory”, thus a memory is disclosed) coupled to the at least one processor, the memory comprising instructions to generate a machine learning decision tree model for a plurality of categorical variables, the instructions, when executed by the at least one processor, to cause the at least one processor to: determine a set of numeric variables for each of the plurality of categorical variables (Gulin, Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list; and (ii) a number of pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value in the respective ordered list.”, thus determining a set of numeric variables for each of the plurality of categorical variables is disclosed because Gulin teaches generating a numeric representation for a categorical feature using occurrence counts and event-outcome counts associated with that categorical feature. Gulin’s categorical feature corresponds to the plurality of categorical variables, while the generated numeric representation based on occurrences and event outcomes corresponds to the set of numeric variables determined for each categorical variable) determine an event rate for each of the plurality of categorical variables (Gulin, Par. [0065], “The MLA translates the value of the categorical feature 122 (i.e. “Pop”) into a numeric feature using a formula: Count = NumberWINs /NumberOCCURENCEs”, & Par. [0103], “generating the numeric representation thereof comprises: (i) using as the number of total occurrences of the at least one preceding training object with the same categorical feature value: a number of total occurrences of the at least one preceding training object having both the first categorical feature value and the second categorical feature value; and (ii) using as the number of the pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value: a number of the pre-determined outcomes of events associated with at least one preceding training object having both the first categorical feature value and the second categorical feature value”, thus determining an event rate for each of the plurality of categorical variables is disclosed because Gulin teaches calculating a count value as a ratio of NumberWINs to NumberOCCURENCEs for categorical features, where the wins correspond to event outcomes and the occurrences correspond to occurrences of the categorical feature values. The ratio generated from event outcomes and occurrences corresponds to an event rate associated with each categorical variable) determine, using training data, […] for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, […] assigning multiple of the plurality of categorical variables to a single node based on the event rate (Gulin, Par. [0093], “the method comprises: accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0082], “For each node of the decision tree, the MLA translates the categorical features (or groups of categorical features, as the case may be) into numeric representation thereof as has been described above”, & Par. [0076], “For example, rather than just looking at the genre of the song, the formula can analyze co-occurrence of the given genre and the given singer (as examples of two categorical features or a group of categorical features that can be associated with a single training object)”, & Par. [0078], “Where both the Number.sub.WINs (F1 and F2) and the Number.sub.OCCURENCEs(F1 and F2) are considering the wins and co-occurrences of the group of features (F1 and F2) values that occur before the current occurrence of the group of features in the ordered list of training objects 102”, thus determining, using training data, for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, and assigning multiple categorical variables to a single node based on the event rate is disclosed because Gulin teaches accessing training objects containing event indicators and categorical features, building a decision tree, and testing which feature and split value to place at each node. Gulin’s training objects correspond to the training data, the categorical features correspond to the plurality of categorical variables, and the decision-tree nodes correspond to the plurality of nodes. Gulin further teaches using groups of categorical features, such as genre and singer/F1 and F2, and calculating wins and occurrences for the grouped features. The wins correspond to event outcomes, the occurrences correspond to feature value occurrences, and the wins/occurrences relationship corresponds to the event rate used to assign multiple categorical variables to a node) and access data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes […] (Gulin, Par. [0093], “accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0221], “The tree model 600 may be referred to as a machine-learning model or a portion of a machine-learning model (e.g., for implementations wherein the machine-learning model relies on multiple tree models). In some instances, the tree model 600 may be referred as a prediction model or a portion of a prediction model (e.g., for implementations wherein the prediction model relies on multiple tree models)”, thus accessing data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes is disclosed because Gulin teaches accessing training objects from a non-transitory computer-readable medium, building a decision tree, and testing which feature and split value to place at the node of each level. Gulin’s training objects correspond to the accessed data, the tree model corresponds to the machine learning decision tree model, and the nodes/levels of the decision tree correspond to the plurality of nodes. The placement of a feature and split value at the node teaches that at least a portion of the nodes are assigned feature/split information during generation of the decision-tree model) Gulin does not explicitly disclose […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] and […] assigned one of the plurality of multi-categorical splits. However, Guillame teaches: […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] (Guillame, Par. [0010], “receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus a plurality of splits wherein the plurality of splits comprises a plurality of multi-categorical splits is disclosed because Guillaume teaches receiving and selecting from a plurality of proposed splits generated from attributes of a training dataset. Guillaume further teaches categorical split conditions for categorical columns and teaches that a super split can refer to a set of splits mapped one-to-one with open leaves of a tree. Guillaume’s plurality of proposed splits corresponds to the plurality of splits, while the super split/set of splits for categorical columns corresponds to the plurality of multi-categorical splits) […] assigned one of the plurality of multi-categorical splits (Guillame, Par. [0010], “The method includes, for each of a plurality of iterations, receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits. The method includes, for each of the plurality of iterations, updating, by the one or more computing devices, a node structure of the decision tree based at least in part on the selected final split”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus assigned one of the plurality of multi-categorical splits is disclosed because Guillaume teaches receiving a plurality of proposed splits, selecting a final split from the plurality of proposed splits, and updating the node structure of the decision tree based on the selected final split. Guillaume further teaches categorical split conditions and teaches that a “super split” can refer to a set of splits mapped one-to-one with open leaves of the tree. Guillaume’s categorical/super splits correspond to the multi-categorical splits, while the updating of the node structure based on the selected final split teaches assigning one of the multi-categorical splits to the decision-tree structure/open leaves) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Gulin’s teaching of event rate based grouped categorical feature processing with Guillaume’s teaching of plurality of proposed splits, super splits, and updating a node structure based on a selected final split. Gulin teaches optimizing which feature and split value to place at each node, while Guillaume teaches an exact distributed split-search approach that provides lower volatile memory usage and smaller space, disk, and network complexity. Therefore, a POSITA would have been motivated to incorporate Guillaume’s split generation and node update techniques into Gulin’s decision tree framework to improve split-selection efficiency and reduce computational resource usage during decision tree training (Guillaume, Par. [0023], “the present disclosure provides an exact distributed algorithm to train Random Forest models as well as other decision forest models without relying on approximating best split search”, &Par. [0036], “only a part of the mapping needs to reside in volatile memory at any instant, which advantageously provides lower volatile memory usage”, & Par. [0038], “The methods described herein stand out from existing exact distributed approaches by a smaller space, disk and network complexity”) Regarding Claim 9, Gulin combined with Guillaume teaches all the limitations of claim 8 as cited above and Guillaume further teaches: determine the multi-categorical splits […] associated with each of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, & Par. [0077], “An example listing is given in Algorithm 7. The iteration over the samples can be trivially parallelized (multithreading over sharding). TABLE-US-00002 Algorithm 7 Find the best supersplits for categorical attribute j and tree p. Nodes are open when they are still subject to splitting - typically nodes are closed when they reach some purity level or when their cardinal is below some threshold. H.sub.h ∈ [1,l] is an empty bi-histogram between the labels and the attribute j for the leaf l for all i in 1,...,n // This loop can be parallelized do h ← sample2node(i) if h is a closed node then continue if candidate feature(j,h,p) is false then continue B ← bag(i,p) // Number of times i is sampled in tree p if B = 0 then continue Add (x.sub.i,j,y.sub.i) weighted by B to H.sub.h end for for all open leaf h do Find best condition using bi-histogram H.sub.h end for”, thus determining the multi-categorical splits associated with each of the plurality of categorical variables is disclosed because Guillaume teaches computing bi-variate histograms between categorical attribute values and label values and identifying optimal split conditions for categorical attributes. Guillaume further teaches finding “best supersplits” for categorical attributes using bi-histograms associated with leaves/nodes of the tree. Guillaume’s categorical attributes correspond to the categorical variables, while the identified optimal split conditions and supersplits correspond to the multi-categorical splits associated with the categorical variables) Guillaume does not explicitly teach […] based on a grouping threshold of the event rate […]. However, Gulin teaches […] based on a grouping threshold of the event rate […] (Gulin, Par. [0148], “The method comprises: generating a range of all possible values of the categorical features; applying a grid to the range to separate the range into region, each region having a boundary; using the boundary as the split value; the generating and the applying being executed before the categorical feature value is translated into the numeric representation thereof”, & Par. [0088], “In a specific embodiment of the present technology, the MLA generates the splits by generating a range of all possible values for the splits (for a given counter having been generated based on the given categorical feature) and applying a pre-determined grid. In some embodiments of the present technology, the range can be between 0 and 1. In other embodiments of the present technology, which is especially pertinent when a coefficient (R.sub.constant) is applied to calculating the values of the counts, the range can be between: (i) the value of the coefficient and (ii) the value of coefficient plus one”, thus based on a grouping threshold of the event rate is disclosed because Gulin teaches generating ranges of values for categorical features/splits, applying a grid to separate the range into regions, and using region boundaries as split values. Gulin further teaches that the split ranges may be generated from counts associated with categorical features. The counts generated from categorical feature occurrences and event outcomes correspond to event rate information, while the region boundaries/split values generated from the grid correspond to grouping thresholds used for determining the splits) Regarding Claim 15, Gulin teaches a non-transitory computer-readable medium (Gulin, Par. [0125], “a non-transitory computer-readable medium”, thus a non-transitory computer-readable medium is disclosed) storing instructions to generate a machine learning decision tree model for a plurality of categorical variables, the instructions configured to cause one or more processors (Gulin, Par. [0125], “a processor”, thus one or more processors is disclosed) of a computing device to determine a set of numeric variables for each of the plurality of categorical variables (Gulin, Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list; and (ii) a number of pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value in the respective ordered list.”, thus determining a set of numeric variables for each of the plurality of categorical variables is disclosed because Gulin teaches generating a numeric representation for a categorical feature using occurrence counts and event-outcome counts associated with that categorical feature. Gulin’s categorical feature corresponds to the plurality of categorical variables, while the generated numeric representation based on occurrences and event outcomes corresponds to the set of numeric variables determined for each categorical variable) determine an event rate for each of the plurality of categorical variables (Gulin, Par. [0065], “The MLA translates the value of the categorical feature 122 (i.e. “Pop”) into a numeric feature using a formula: Count = NumberWINs /NumberOCCURENCEs”, & Par. [0103], “generating the numeric representation thereof comprises: (i) using as the number of total occurrences of the at least one preceding training object with the same categorical feature value: a number of total occurrences of the at least one preceding training object having both the first categorical feature value and the second categorical feature value; and (ii) using as the number of the pre-determined outcomes of events associated with at least one preceding training object having the same categorical feature value: a number of the pre-determined outcomes of events associated with at least one preceding training object having both the first categorical feature value and the second categorical feature value”, thus determining an event rate for each of the plurality of categorical variables is disclosed because Gulin teaches calculating a count value as a ratio of NumberWINs to NumberOCCURENCEs for categorical features, where the wins correspond to event outcomes and the occurrences correspond to occurrences of the categorical feature values. The ratio generated from event outcomes and occurrences corresponds to an event rate associated with each categorical variable) determine, using training data, […] for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, […] assigning multiple of the plurality of categorical variables to a single node based on the event rate (Gulin, Par. [0093], “the method comprises: accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0082], “For each node of the decision tree, the MLA translates the categorical features (or groups of categorical features, as the case may be) into numeric representation thereof as has been described above”, & Par. [0076], “For example, rather than just looking at the genre of the song, the formula can analyze co-occurrence of the given genre and the given singer (as examples of two categorical features or a group of categorical features that can be associated with a single training object)”, & Par. [0078], “Where both the Number.sub.WINs (F1 and F2) and the Number.sub.OCCURENCEs(F1 and F2) are considering the wins and co-occurrences of the group of features (F1 and F2) values that occur before the current occurrence of the group of features in the ordered list of training objects 102”, thus determining, using training data, for assigning the plurality of categorical variables to a plurality of nodes of the machine learning decision tree model, and assigning multiple categorical variables to a single node based on the event rate is disclosed because Gulin teaches accessing training objects containing event indicators and categorical features, building a decision tree, and testing which feature and split value to place at each node. Gulin’s training objects correspond to the training data, the categorical features correspond to the plurality of categorical variables, and the decision-tree nodes correspond to the plurality of nodes. Gulin further teaches using groups of categorical features, such as genre and singer/F1 and F2, and calculating wins and occurrences for the grouped features. The wins correspond to event outcomes, the occurrences correspond to feature value occurrences, and the wins/occurrences relationship corresponds to the event rate used to assign multiple categorical variables to a node) and access data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes […] (Gulin, Par. [0093], “accessing, from a non-transitory computer-readable medium of the machine learning system, a set of training objects, each training object of the set of training object containing a document and an event indicator associated with the document, each document being associated with a categorical feature”, & Par. [0092], “When the MLA builds a decision tree at a particular iteration of the decision tree model building, for each level, the MLA tests and optimizes the best of: which feature to place at the node of the level and which split value (out of all possible pre-defined values) to place at the node”, & Par. [0221], “The tree model 600 may be referred to as a machine-learning model or a portion of a machine-learning model (e.g., for implementations wherein the machine-learning model relies on multiple tree models). In some instances, the tree model 600 may be referred as a prediction model or a portion of a prediction model (e.g., for implementations wherein the prediction model relies on multiple tree models)”, thus accessing data to generate the machine learning decision tree model comprising the plurality of nodes, at least a portion of the plurality of nodes is disclosed because Gulin teaches accessing training objects from a non-transitory computer-readable medium, building a decision tree, and testing which feature and split value to place at the node of each level. Gulin’s training objects correspond to the accessed data, the tree model corresponds to the machine learning decision tree model, and the nodes/levels of the decision tree correspond to the plurality of nodes. The placement of a feature and split value at the node teaches that at least a portion of the nodes are assigned feature/split information during generation of the decision-tree model) Gulin does not explicitly disclose […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] and […] assigned one of the plurality of multi-categorical splits. However, Guillame teaches: […] a plurality of splits […] wherein the plurality of splits comprises a plurality of multi-categorical splits […] (Guillame, Par. [0010], “receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus a plurality of splits wherein the plurality of splits comprises a plurality of multi-categorical splits is disclosed because Guillaume teaches receiving and selecting from a plurality of proposed splits generated from attributes of a training dataset. Guillaume further teaches categorical split conditions for categorical columns and teaches that a super split can refer to a set of splits mapped one-to-one with open leaves of a tree. Guillaume’s plurality of proposed splits corresponds to the plurality of splits, while the super split/set of splits for categorical columns corresponds to the plurality of multi-categorical splits) […] assigned one of the plurality of multi-categorical splits (Guillame, Par. [0010], “The method includes, for each of a plurality of iterations, receiving, by the one or more computing devices, a plurality of proposed splits from a plurality of splitters. The plurality of proposed splits is respectively generated based on a plurality of attributes of a training dataset. The method includes, for each of the plurality of iterations, selecting, by the one or more computing devices, a final split from the plurality of proposed splits. The method includes, for each of the plurality of iterations, updating, by the one or more computing devices, a node structure of the decision tree based at least in part on the selected final split”, & Par. [0072], “For categorical columns, the condition is of the form x.sub.i,j ∈ C with C ⊆ S.sub.j and S.sub.j the support of column j. In case of attribute sampling (e.g. RF), only a random subset of attributes are considered. The super split can refer to a set of splits mapped one-to-one with the open leaves at a given depth of a tree”, thus assigned one of the plurality of multi-categorical splits is disclosed because Guillaume teaches receiving a plurality of proposed splits, selecting a final split from the plurality of proposed splits, and updating the node structure of the decision tree based on the selected final split. Guillaume further teaches categorical split conditions and teaches that a “super split” can refer to a set of splits mapped one-to-one with open leaves of the tree. Guillaume’s categorical/super splits correspond to the multi-categorical splits, while the updating of the node structure based on the selected final split teaches assigning one of the multi-categorical splits to the decision-tree structure/open leaves) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Gulin’s teaching of event rate based grouped categorical feature processing with Guillaume’s teaching of plurality of proposed splits, super splits, and updating a node structure based on a selected final split. Gulin teaches optimizing which feature and split value to place at each node, while Guillaume teaches an exact distributed split-search approach that provides lower volatile memory usage and smaller space, disk, and network complexity. Therefore, a POSITA would have been motivated to incorporate Guillaume’s split generation and node update techniques into Gulin’s decision tree framework to improve split-selection efficiency and reduce computational resource usage during decision tree training (Guillaume, Par. [0023], “the present disclosure provides an exact distributed algorithm to train Random Forest models as well as other decision forest models without relying on approximating best split search”, &Par. [0036], “only a part of the mapping needs to reside in volatile memory at any instant, which advantageously provides lower volatile memory usage”, & Par. [0038], “The methods described herein stand out from existing exact distributed approaches by a smaller space, disk and network complexity”) Regarding Claim 16, Gulin combined with Guillaume teaches all the limitations of claim 15 as cited above and Guillaume further teaches: determine the multi-categorical splits […] associated with each of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, & Par. [0077], “An example listing is given in Algorithm 7. The iteration over the samples can be trivially parallelized (multithreading over sharding). TABLE-US-00002 Algorithm 7 Find the best supersplits for categorical attribute j and tree p. Nodes are open when they are still subject to splitting - typically nodes are closed when they reach some purity level or when their cardinal is below some threshold. H.sub.h ∈ [1,l] is an empty bi-histogram between the labels and the attribute j for the leaf l for all i in 1,...,n // This loop can be parallelized do h ← sample2node(i) if h is a closed node then continue if candidate feature(j,h,p) is false then continue B ← bag(i,p) // Number of times i is sampled in tree p if B = 0 then continue Add (x.sub.i,j,y.sub.i) weighted by B to H.sub.h end for for all open leaf h do Find best condition using bi-histogram H.sub.h end for”, thus determining the multi-categorical splits associated with each of the plurality of categorical variables is disclosed because Guillaume teaches computing bi-variate histograms between categorical attribute values and label values and identifying optimal split conditions for categorical attributes. Guillaume further teaches finding “best supersplits” for categorical attributes using bi-histograms associated with leaves/nodes of the tree. Guillaume’s categorical attributes correspond to the categorical variables, while the identified optimal split conditions and supersplits correspond to the multi-categorical splits associated with the categorical variables) Guillaume does not explicitly teach […] based on a grouping threshold of the event rate […]. However, Gulin teaches […] based on a grouping threshold of the event rate […] (Gulin, Par. [0148], “The method comprises: generating a range of all possible values of the categorical features; applying a grid to the range to separate the range into region, each region having a boundary; using the boundary as the split value; the generating and the applying being executed before the categorical feature value is translated into the numeric representation thereof”, & Par. [0088], “In a specific embodiment of the present technology, the MLA generates the splits by generating a range of all possible values for the splits (for a given counter having been generated based on the given categorical feature) and applying a pre-determined grid. In some embodiments of the present technology, the range can be between 0 and 1. In other embodiments of the present technology, which is especially pertinent when a coefficient (R.sub.constant) is applied to calculating the values of the counts, the range can be between: (i) the value of the coefficient and (ii) the value of coefficient plus one”, thus based on a grouping threshold of the event rate is disclosed because Gulin teaches generating ranges of values for categorical features/splits, applying a grid to separate the range into regions, and using region boundaries as split values. Gulin further teaches that the split ranges may be generated from counts associated with categorical features. The counts generated from categorical feature occurrences and event outcomes correspond to event rate information, while the region boundaries/split values generated from the grid correspond to grouping thresholds used for determining the splits) 07-21-aia AIA Claim s 3, 10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Gulin et al. (hereafter Gulin) (US 2019164085), in view of Guillame et al. (hereinafter Guillame) (US 2019251468) and further in view of Pingenot et al. (hereinafter Pingenot) (US 9633311). Regarding Claim 3, Gulin combined with Guillaume teaches all the limitations of claim 1 as cited above and Gulin further teaches: […] of the machine learning decision tree model (Gulin, Abstract, “There is disclosed a method of and a system for training and using a Machine Learning Algorithm (MLA), the MLA using a decision tree model having a decision tree”, thus a machine learning decision tree model is disclosed because Gulin teaches a Machine Learning Algorithm (MLA) using a decision tree model having a decision tree) Guillame further teaches […] the multi-categorical splits […] (Guillame, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the multi-categorical splits are disclosed because Guillaume teaches determining optimal split conditions for categorical attributes using bi-variate histograms between attribute values and label values. Guillaume’s identified optimal split conditions for categorical attributes correspond to the multi-categorical splits) Gulin combined with Guillame does not explicitly teach determining a specified depth […] and determining […] based on the specified depth. However, Pingenot teaches determining a specified depth […] (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc”, thus determining a specified depth is disclosed because Pingenot teaches that recursive partitioning of the decision tree continues until a stopping condition is met, including when a specific depth of the decision tree is reached) Pingenot further teaches and determining […] based on the specified depth (Pingenot, Par. [0035], “In an operation 206, a fourth indicator of one or more decision tree parameters is received. For example, the fourth indicator may indicate a maximum depth of the decision tree, a threshold worth value, a minimum number of observations per leaf node, etc.”, thus determining based on the specified depth is disclosed because Pingenot teaches receiving a decision-tree parameter indicating a maximum depth of the decision tree, where the maximum depth parameter is used during decision-tree generation/training and therefore influences subsequent split determination and tree construction operations) It would have been obvious to combine Gulin and Guillaume with Pingenot because Pingenot teaches that decision tree learning algorithms select variables and splits that best divide the data and further teaches that different algorithms may use different metrics for determining the “best” split. Gulin teaches event rate based categorical feature processing and grouped categorical feature split generation, while Guillaume teaches plurality of proposed splits, super splits, and updating a node structure based on selected splits. Pingenot further teaches controlling recursive partitioning using a specified decision-tree depth. Therefore, a POSITA would have been motivated to incorporate Pingenot’s specified-depth stopping condition into the Gulin/Guillaume decision tree framework to improve control of recursive partitioning and split generation during decision tree training (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc. Algorithms for constructing decision trees (decision tree learning) usually work top-down by choosing a variable at each step that best splits the set of items. Different algorithms use different metrics for measuring the “best” split”) Regarding Claim 10, Gulin combined with Guillaume teaches all the limitations of claim 8 as cited above and Gulin further teaches: […] of the machine learning decision tree model (Gulin, Abstract, “There is disclosed a method of and a system for training and using a Machine Learning Algorithm (MLA), the MLA using a decision tree model having a decision tree”, thus a machine learning decision tree model is disclosed because Gulin teaches a Machine Learning Algorithm (MLA) using a decision tree model having a decision tree) Guillame further teaches […] the multi-categorical splits […] (Guillame, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the multi-categorical splits are disclosed because Guillaume teaches determining optimal split conditions for categorical attributes using bi-variate histograms between attribute values and label values. Guillaume’s identified optimal split conditions for categorical attributes correspond to the multi-categorical splits) Gulin combined with Guillame does not explicitly teach determine a specified depth […] and determine […] based on the specified depth. However, Pingenot teaches determine a specified depth […] (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc”, thus determining a specified depth is disclosed because Pingenot teaches that recursive partitioning of the decision tree continues until a stopping condition is met, including when a specific depth of the decision tree is reached) Pingenot further teaches determine […] based on the specified depth (Pingenot, Par. [0035], “In an operation 206, a fourth indicator of one or more decision tree parameters is received. For example, the fourth indicator may indicate a maximum depth of the decision tree, a threshold worth value, a minimum number of observations per leaf node, etc.”, thus determining based on the specified depth is disclosed because Pingenot teaches receiving a decision-tree parameter indicating a maximum depth of the decision tree, where the maximum depth parameter is used during decision-tree generation/training and therefore influences subsequent split determination and tree construction operations) It would have been obvious to combine Gulin and Guillaume with Pingenot because Pingenot teaches that decision tree learning algorithms select variables and splits that best divide the data and further teaches that different algorithms may use different metrics for determining the “best” split. Gulin teaches event rate based categorical feature processing and grouped categorical feature split generation, while Guillaume teaches plurality of proposed splits, super splits, and updating a node structure based on selected splits. Pingenot further teaches controlling recursive partitioning using a specified decision-tree depth. Therefore, a POSITA would have been motivated to incorporate Pingenot’s specified-depth stopping condition into the Gulin/Guillaume decision tree framework to improve control of recursive partitioning and split generation during decision tree training (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc. Algorithms for constructing decision trees (decision tree learning) usually work top-down by choosing a variable at each step that best splits the set of items. Different algorithms use different metrics for measuring the “best” split”) Regarding Claim 17, Gulin combined with Guillaume teaches all the limitations of claim 15 as cited above and Gulin further teaches: […] of the machine learning decision tree model (Gulin, Abstract, “There is disclosed a method of and a system for training and using a Machine Learning Algorithm (MLA), the MLA using a decision tree model having a decision tree”, thus a machine learning decision tree model is disclosed because Gulin teaches a Machine Learning Algorithm (MLA) using a decision tree model having a decision tree) Guillame further teaches […] the multi-categorical splits […] (Guillame, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the multi-categorical splits are disclosed because Guillaume teaches determining optimal split conditions for categorical attributes using bi-variate histograms between attribute values and label values. Guillaume’s identified optimal split conditions for categorical attributes correspond to the multi-categorical splits) Gulin combined with Guillame does not explicitly teach determine a specified depth […] and determine […] based on the specified depth. However, Pingenot teaches determine a specified depth […] (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc”, thus determining a specified depth is disclosed because Pingenot teaches that recursive partitioning of the decision tree continues until a stopping condition is met, including when a specific depth of the decision tree is reached) Pingenot further teaches determine […] based on the specified depth (Pingenot, Par. [0035], “In an operation 206, a fourth indicator of one or more decision tree parameters is received. For example, the fourth indicator may indicate a maximum depth of the decision tree, a threshold worth value, a minimum number of observations per leaf node, etc.”, thus determining based on the specified depth is disclosed because Pingenot teaches receiving a decision-tree parameter indicating a maximum depth of the decision tree, where the maximum depth parameter is used during decision-tree generation/training and therefore influences subsequent split determination and tree construction operations) It would have been obvious to combine Gulin and Guillaume with Pingenot because Pingenot teaches that decision tree learning algorithms select variables and splits that best divide the data and further teaches that different algorithms may use different metrics for determining the “best” split. Gulin teaches event rate based categorical feature processing and grouped categorical feature split generation, while Guillaume teaches plurality of proposed splits, super splits, and updating a node structure based on selected splits. Pingenot further teaches controlling recursive partitioning using a specified decision-tree depth. Therefore, a POSITA would have been motivated to incorporate Pingenot’s specified-depth stopping condition into the Gulin/Guillaume decision tree framework to improve control of recursive partitioning and split generation during decision tree training (Pingenot, Par. [0015], “A tree can be “learned” by splitting the source dataset into two or more subsets based on a test of the attribute value of a specific variable. This process is repeated on each derived subset in a recursive manner called recursive partitioning. The recursion is completed when the subset at a node has all the same value of the target variable, when splitting no longer increases an estimated value of the prediction, when a specific depth of the decision tree is reached, etc. Algorithms for constructing decision trees (decision tree learning) usually work top down by choosing a variable at each step that best splits the set of items. Different algorithms use different metrics for measuring the “best” split”) 07-21-aia AIA Claim s 4-5, 11-12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Gulin et al. (hereafter Gulin) (US 2019164085), in view of Guillame et al. (hereinafter Guillame) (US 2019251468) and further in view of Kim et al. (hereinafter Kim), a non-patent literature reference titled “A hybrid decision tree algorithm for mixed numeric and categorical data in regression analysis”). Regarding Claim 4, Gulin combined with Guillaume teaches all the limitations of claim 1 as cited above and Gulin further teaches: […] each of the plurality of categorical variables […] of the plurality of categorical variables (Gulin, Par. [0041], “A typical prior art approach, is to convert the categorical feature into a numeric representation thereof and processing the numeric representation of the categorical feature. There are various known approaches to converting categorical features into numeric representation: categorical encoding, numeric encoding, one-hot encoding, binary encoding, etc”, & Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list;”, thus each of the plurality of categorical variables of the plurality of categorical variables is disclosed because Gulin teaches converting categorical features into numeric representations using encoding techniques such as categorical encoding, one hot encoding, and binary encoding, and further teaches processing a given categorical feature by generating a numeric representation thereof. Gulin’s categorical features correspond to the plurality of categorical variables, while the generated numeric representations correspond to processing each categorical variable of the plurality of categorical variables) Gulin combined with Guillame does not explicitly teach replacing […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...]. However, Kim teaches replacing […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...] (Kim, Page 1 – Section 1, “One possible conversion of categorical variable data to numeric data is the use of various coding systems such as dummy coding, effects coding, and contrast coding to manage the presence of categorical predictors”, & Page 2 – Section 2.1, “Nominal variables are converted to quantitative data through dummy coding, effects coding, contrast coding and so on”, & Page 4 – Section 4, “The maximum cardinality is defined as the number of the total variables in the final data after dummy variable creation”, thus replacing […categorical variables…] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the […categorical variables…] is disclosed because Kim teaches converting categorical variable data into numeric data using dummy coding and teaches creation of dummy variables in the final dataset representation. Kim’s dummy coding replaces the […categorical variables…] with multiple dummy variables, where the dummy variables represent different categorical values associated with the […categorical variables…]) It would have been obvious to combine Gulin and Guillaume with Kim because Gulin and Guillaume teach generating and processing categorical feature-based decision tree splits, while Kim teaches converting categorical variables into dummy coded variables for improved regression model performance. Kim further teaches that the dummy-coded categorical variables usually improved the performance and that the proposed hybrid method outperformed comparison models across datasets. Therefore, a POSITA would have been motivated to incorporate Kim’s dummy variable conversion techniques into the Gulin/Guillaume decision tree framework to improve processing of categorical variables and improve machine-learning model performance during decision tree training and regression analysis (Kim, Page 4 – Section 4.3, “Based on Table 2 , regardless of the datasets, the proposed method outperformed M1 and M2. The improvements in the hybrid method in terms of MSE depended on the regression algorithm combined with the decision tree. Comparing M1 and M2, the dummy-coded categorical variables usually improved the performance”) Regarding Claim 5, Gulin and Guillaume combined with Kim teaches all the limitations of claim 4 as cited above and Guillaume further teaches: […] the plurality of splits for the plurality of categorical variables […] (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, & Par. [0077], “An example listing is given in Algorithm 7. The iteration over the samples can be trivially parallelized (multithreading over sharding). TABLE-US-00002 Algorithm 7 Find the best supersplits for categorical attribute j and tree p. Nodes are open when they are still subject to splitting - typically nodes are closed when they reach some purity level or when their cardinal is below some threshold. H.sub.h ∈ [1,l] is an empty bi-histogram between the labels and the attribute j for the leaf l for all i in 1,...,n // This loop can be parallelized do h ← sample2node(i) if h is a closed node then continue if candidate feature(j,h,p) is false then continue B ← bag(i,p) // Number of times i is sampled in tree p if B = 0 then continue Add (x.sub.i,j,y.sub.i) weighted by B to H.sub.h end for for all open leaf h do Find best condition using bi-histogram H.sub.h end for”, thus the plurality of splits for the plurality of categorical variables is disclosed because Guillaume teaches identifying optimal or approximate splits for categorical attributes using bi-variate histograms between attribute values and label values. Guillaume further teaches finding best super splits for categorical attribute j and tree p, where a best condition is found using the bi-histogram for each open leaf. Guillaume’s categorical attributes correspond to the plurality of categorical variables, while the identified optimal splits/supersplits correspond to the plurality of splits determined for those categorical variables) Gulin combined with Guillaume does not explicitly teach separately determining […] for each of the plurality of dummy variables. However, Kim teaches separately determining […] for each of the plurality of dummy variables (Kim, Page 2 – Section 2.1, “Nominal variables are converted to quantitative data through dummy coding, effects coding, contrast coding and so on”, & Page 3 – Section 3, “The proposed algorithm trains explained variations in a target variable separately by categorical and numeric variables. The two parts of the variations in the target are estimated in series. The combinations of values of several categorical variables decide their effect on the target by utilizing the decision tree. Then, the remaining part, unexplained by categorical variables, is estimated by numeric variables through one of the many regression algorithms”, thus separately determining for each of the plurality of dummy variables is disclosed because Kim teaches converting nominal variables into quantitative data through dummy coding and further teaches separately training/estimating variations associated with categorical and numeric variables. Kim’s dummy coding corresponds to the plurality of dummy variables, while the separate estimation/training of the variable effects corresponds to separately determining information for each of the plurality of dummy variables) Regarding Claim 11, Gulin combined with Guillaume teaches all the limitations of claim 8 as cited above and Gulin further teaches: […] each of the plurality of categorical variables […] of the plurality of categorical variables (Gulin, Par. [0041], “A typical prior art approach, is to convert the categorical feature into a numeric representation thereof and processing the numeric representation of the categorical feature. There are various known approaches to converting categorical features into numeric representation: categorical encoding, numeric encoding, one-hot encoding, binary encoding, etc”, & Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list;”, thus each of the plurality of categorical variables of the plurality of categorical variables is disclosed because Gulin teaches converting categorical features into numeric representations using encoding techniques such as categorical encoding, one hot encoding, and binary encoding, and further teaches processing a given categorical feature by generating a numeric representation thereof. Gulin’s categorical features correspond to the plurality of categorical variables, while the generated numeric representations correspond to processing each categorical variable of the plurality of categorical variables) Gulin combined with Guillame does not explicitly teach replace […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...]. However, Kim teaches replace […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...] (Kim, Page 1 – Section 1, “One possible conversion of categorical variable data to numeric data is the use of various coding systems such as dummy coding, effects coding, and contrast coding to manage the presence of categorical predictors”, & Page 2 – Section 2.1, “Nominal variables are converted to quantitative data through dummy coding, effects coding, contrast coding and so on”, & Page 4 – Section 4, “The maximum cardinality is defined as the number of the total variables in the final data after dummy variable creation”, thus replacing […categorical variables…] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the […categorical variables…] is disclosed because Kim teaches converting categorical variable data into numeric data using dummy coding and teaches creation of dummy variables in the final dataset representation. Kim’s dummy coding replaces the […categorical variables…] with multiple dummy variables, where the dummy variables represent different categorical values associated with the […categorical variables…]) It would have been obvious to combine Gulin and Guillaume with Kim because Gulin and Guillaume teach generating and processing categorical feature-based decision tree splits, while Kim teaches converting categorical variables into dummy coded variables for improved regression model performance. Kim further teaches that the dummy-coded categorical variables usually improved the performance and that the proposed hybrid method outperformed comparison models across datasets. Therefore, a POSITA would have been motivated to incorporate Kim’s dummy variable conversion techniques into the Gulin/Guillaume decision tree framework to improve processing of categorical variables and improve machine-learning model performance during decision tree training and regression analysis (Kim, Page 4 – Section 4.3, “Based on Table 2 , regardless of the datasets, the proposed method outperformed M1 and M2. The improvements in the hybrid method in terms of MSE depended on the regression algorithm combined with the decision tree. Comparing M1 and M2, the dummy-coded categorical variables usually improved the performance”) Regarding Claim 12, Gulin and Guillaume combined with Kim teaches all the limitations of claim 11 as cited above and Guillaume further teaches: […] the plurality of splits for the plurality of categorical variables […] (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, & Par. [0077], “An example listing is given in Algorithm 7. The iteration over the samples can be trivially parallelized (multithreading over sharding). TABLE-US-00002 Algorithm 7 Find the best supersplits for categorical attribute j and tree p. Nodes are open when they are still subject to splitting - typically nodes are closed when they reach some purity level or when their cardinal is below some threshold. H.sub.h ∈ [1,l] is an empty bi-histogram between the labels and the attribute j for the leaf l for all i in 1,...,n // This loop can be parallelized do h ← sample2node(i) if h is a closed node then continue if candidate feature(j,h,p) is false then continue B ← bag(i,p) // Number of times i is sampled in tree p if B = 0 then continue Add (x.sub.i,j,y.sub.i) weighted by B to H.sub.h end for for all open leaf h do Find best condition using bi-histogram H.sub.h end for”, thus the plurality of splits for the plurality of categorical variables is disclosed because Guillaume teaches identifying optimal or approximate splits for categorical attributes using bi-variate histograms between attribute values and label values. Guillaume further teaches finding best super splits for categorical attribute j and tree p, where a best condition is found using the bi-histogram for each open leaf. Guillaume’s categorical attributes correspond to the plurality of categorical variables, while the identified optimal splits/supersplits correspond to the plurality of splits determined for those categorical variables) Gulin combined with Guillaume does not explicitly teach separately determine […] for each of the plurality of dummy variables. However, Kim teaches separately determine […] for each of the plurality of dummy variables (Kim, Page 2 – Section 2.1, “Nominal variables are converted to quantitative data through dummy coding, effects coding, contrast coding and so on”, & Page 3 – Section 3, “The proposed algorithm trains explained variations in a target variable separately by categorical and numeric variables. The two parts of the variations in the target are estimated in series. The combinations of values of several categorical variables decide their effect on the target by utilizing the decision tree. Then, the remaining part, unexplained by categorical variables, is estimated by numeric variables through one of the many regression algorithms”, thus separately determining for each of the plurality of dummy variables is disclosed because Kim teaches converting nominal variables into quantitative data through dummy coding and further teaches separately training/estimating variations associated with categorical and numeric variables. Kim’s dummy coding corresponds to the plurality of dummy variables, while the separate estimation/training of the variable effects corresponds to separately determining information for each of the plurality of dummy variables) Regarding Claim 18, Gulin combined with Guillaume teaches all the limitations of claim 15 as cited above and Gulin further teaches: […] each of the plurality of categorical variables […] of the plurality of categorical variables (Gulin, Par. [0041], “A typical prior art approach, is to convert the categorical feature into a numeric representation thereof and processing the numeric representation of the categorical feature. There are various known approaches to converting categorical features into numeric representation: categorical encoding, numeric encoding, one-hot encoding, binary encoding, etc”, & Par. [0093], “when processing a given categorical feature using the decision tree structure, the given categorical feature associated with a given training object, the given training object having at least one preceding training object in the ordered list of training objects, generating a numeric representation thereof, the generating based on: (i) a number of total occurrences of the at least one preceding training object with a same categorical feature value in the respective ordered list;”, thus each of the plurality of categorical variables of the plurality of categorical variables is disclosed because Gulin teaches converting categorical features into numeric representations using encoding techniques such as categorical encoding, one hot encoding, and binary encoding, and further teaches processing a given categorical feature by generating a numeric representation thereof. Gulin’s categorical features correspond to the plurality of categorical variables, while the generated numeric representations correspond to processing each categorical variable of the plurality of categorical variables) Gulin combined with Guillame does not explicitly teach replace […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...]. However, Kim teaches replace […] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion [...] (Kim, Page 1 – Section 1, “One possible conversion of categorical variable data to numeric data is the use of various coding systems such as dummy coding, effects coding, and contrast coding to manage the presence of categorical predictors”, & Page 2 – Section 2.1, “Nominal variables are converted to quantitative data through dummy coding, effects coding, contrast coding and so on”, & Page 4 – Section 4, “The maximum cardinality is defined as the number of the total variables in the final data after dummy variable creation”, thus replacing […categorical variables…] with a set of a plurality of dummy variables, wherein each of the plurality of dummy variables is associated with a different value of at least a portion of the […categorical variables…] is disclosed because Kim teaches converting categorical variable data into numeric data using dummy coding and teaches creation of dummy variables in the final dataset representation. Kim’s dummy coding replaces the […categorical variables…] with multiple dummy variables, where the dummy variables represent different categorical values associated with the […categorical variables…]) It would have been obvious to combine Gulin and Guillaume with Kim because Gulin and Guillaume teach generating and processing categorical feature-based decision tree splits, while Kim teaches converting categorical variables into dummy coded variables for improved regression model performance. Kim further teaches that the dummy-coded categorical variables usually improved the performance and that the proposed hybrid method outperformed comparison models across datasets. Therefore, a POSITA would have been motivated to incorporate Kim’s dummy variable conversion techniques into the Gulin/Guillaume decision tree framework to improve processing of categorical variables and improve machine-learning model performance during decision tree training and regression analysis (Kim, Page 4 – Section 4.3, “Based on Table 2 , regardless of the datasets, the proposed method outperformed M1 and M2. The improvements in the hybrid method in terms of MSE depended on the regression algorithm combined with the decision tree. Comparing M1 and M2, the dummy-coded categorical variables usually improved the performance”) 07-21-aia AIA Claim s 6-7, 13-14, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gulin et al. (hereafter Gulin) (US 2019164085), in view of Guillame et al. (hereinafter Guillame) (US 2019251468) and further in view of Pearson et al. (hereinafter Pearson), a non-patent literature reference titled “The problem of disguised missing data”). Regarding Claim 6, Gulin combined with Guillaume teaches all the limitations of claim 1 as cited above and Guillaume further teaches: […] of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the plurality of categorical variables is disclosed because Guillaume teaches determining optimal split conditions for a categorical attribute using attribute values and label values. Guillaume’s categorical attribute corresponds to one of the plurality of categorical variables, and the use of categorical attributes for split determination supports processing the plurality of categorical variables) Gulin combined with Guillaume does not explicitly teach determining missing values […] and determining whether the missing values are informative based on a statistical significance of a category […]. However, Pearson teaches determining missing values […] (Pearson, Page 2 – Section 2, “Key to the use of any of the missing data treatment strategies just described is the recognition that certain data values are missing. The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 5 – Section 3.4, “The S-plus procedure tree() also provides another option for handling missing data (na.action=na.tree.replace.all),which converts all incomplete variables into factor (i.e., categorical)data types,with missing values all assigned to a special “missing” category”, thus determining missing values is disclosed because Pearson teaches recognizing when data values are missing and further teaches handling incomplete variables by assigning missing values to a special “missing” category. Pearson’s recognition of missing/disguised missing data and identification of incomplete variables correspond to determining missing values) Pearson further teaches determining whether the missing values are informative based on a statistical significance of a category […] (Pearson, Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus determining whether the missing values are informative based on a statistical significance of a category is disclosed because Pearson teaches comparing analysis results when missing/disguised missing values are included versus omitted and determining statistical significance using a t-statistic and p-value. Pearson further teaches that patients with missing serum insulin values differ from patients without missing values, raising the possibility of nonignorable missing data. Pearson’s diabetic/nondiabetic groups and missing/non-missing groups correspond to categories, while the p-value and statistically significant difference correspond to determining whether the missing values are informative based on statistical significance) It would have been obvious to combine Gulin and Guillaume with Pearson because Pearson teaches that disguised missing data can be misinterpreted as valid data and can significantly distort statistical analysis results, while Gulin and Guillaume teach determining split conditions and processing categorical variables for decision tree generation. Pearson further teaches identifying missing values, assigning incomplete variables to categorical missing groups, and using imputation techniques when missing data is informative. Therefore, a POSITA would have been motivated to incorporate Pearson’s missing value detection and handling techniques into the categorical split processing framework of Gulin and Guillaume to improve the reliability and accuracy of categorical split generation and decision tree training when incomplete or disguised missing data is present (Pearson, Page 2 – Section 2, “The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”) . Regarding Claim 7, Gulin and Guillaume combined with Pearson teaches all the limitations of claim 6 as cited above and Pearson further teaches: imputing the missing values responsive to the missing values being informative (Pearson, Page 2 – Section 1.3, “For larger fractions of missing data, or in other cases where deletion strategies are deemed undesirable, one common alternative is imputation, where missing data values are estimated on the basis of those that are available”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus imputing the missing values responsive to the missing values being informative is disclosed because Pearson teaches imputing missing data values by estimating them based on available data and further teaches that differences between groups with missing and non-missing values may be important, raising the possibility of nonignorable missing data. Pearson’s imputation of missing data values corresponds to imputing the missing values, while the determination that the missing values may be important/nonignorable corresponds to the missing values being informative) Regarding Claim 13, Gulin combined with Guillaume teaches all the limitations of claim 8 as cited above and Guillaume further teaches: […] of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the plurality of categorical variables is disclosed because Guillaume teaches determining optimal split conditions for a categorical attribute using attribute values and label values. Guillaume’s categorical attribute corresponds to one of the plurality of categorical variables, and the use of categorical attributes for split determination supports processing the plurality of categorical variables) Gulin combined with Guillaume does not explicitly teach determine missing values […] and determine whether the missing values are informative based on a statistical significance of a category […]. However, Pearson teaches determine missing values […] (Pearson, Page 2 – Section 2, “Key to the use of any of the missing data treatment strategies just described is the recognition that certain data values are missing. The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 5 – Section 3.4, “The S-plus procedure tree() also provides another option for handling missing data (na.action=na.tree.replace.all),which converts all incomplete variables into factor (i.e., categorical)data types,with missing values all assigned to a special “missing” category”, thus determining missing values is disclosed because Pearson teaches recognizing when data values are missing and further teaches handling incomplete variables by assigning missing values to a special “missing” category. Pearson’s recognition of missing/disguised missing data and identification of incomplete variables correspond to determining missing values) Pearson further teaches determine whether the missing values are informative based on a statistical significance of a category […] (Pearson, Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus determining whether the missing values are informative based on a statistical significance of a category is disclosed because Pearson teaches comparing analysis results when missing/disguised missing values are included versus omitted and determining statistical significance using a t-statistic and p-value. Pearson further teaches that patients with missing serum insulin values differ from patients without missing values, raising the possibility of nonignorable missing data. Pearson’s diabetic/nondiabetic groups and missing/non-missing groups correspond to categories, while the p-value and statistically significant difference correspond to determining whether the missing values are informative based on statistical significance) It would have been obvious to combine Gulin and Guillaume with Pearson because Pearson teaches that disguised missing data can be misinterpreted as valid data and can significantly distort statistical analysis results, while Gulin and Guillaume teach determining split conditions and processing categorical variables for decision tree generation. Pearson further teaches identifying missing values, assigning incomplete variables to categorical missing groups, and using imputation techniques when missing data is informative. Therefore, a POSITA would have been motivated to incorporate Pearson’s missing value detection and handling techniques into the categorical split processing framework of Gulin and Guillaume to improve the reliability and accuracy of categorical split generation and decision tree training when incomplete or disguised missing data is present (Pearson, Page 2 – Section 2, “The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”) . Regarding Claim 14, Gulin and Guillaume combined with Pearson teaches all the limitations of claim 13 as cited above and Pearson further teaches: impute the missing values responsive to the missing values being informative (Pearson, Page 2 – Section 1.3, “For larger fractions of missing data, or in other cases where deletion strategies are deemed undesirable, one common alternative is imputation, where missing data values are estimated on the basis of those that are available”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus imputing the missing values responsive to the missing values being informative is disclosed because Pearson teaches imputing missing data values by estimating them based on available data and further teaches that differences between groups with missing and non-missing values may be important, raising the possibility of nonignorable missing data. Pearson’s imputation of missing data values corresponds to imputing the missing values, while the determination that the missing values may be important/nonignorable corresponds to the missing values being informative) Regarding Claim 19, Gulin combined with Guillaume teaches all the limitations of claim 15 as cited above and Guillaume further teaches: […] of the plurality of categorical variables (Guillaume, Par. [0075], “Estimating the best condition for a categorical attribute j and in leaf h can include computing the bi-variate histogram between the attribute values and the label values for all the samples in h. The optimal (in case of binary labels) or approximate (in case of multiclass labels) split can then be identified using any number of techniques”, thus the plurality of categorical variables is disclosed because Guillaume teaches determining optimal split conditions for a categorical attribute using attribute values and label values. Guillaume’s categorical attribute corresponds to one of the plurality of categorical variables, and the use of categorical attributes for split determination supports processing the plurality of categorical variables) Gulin combined with Guillaume does not explicitly teach determine missing values […] and determine whether the missing values are informative based on a statistical significance of a category […]. However, Pearson teaches determine missing values […] (Pearson, Page 2 – Section 2, “Key to the use of any of the missing data treatment strategies just described is the recognition that certain data values are missing. The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 5 – Section 3.4, “The S-plus procedure tree() also provides another option for handling missing data (na.action=na.tree.replace.all),which converts all incomplete variables into factor (i.e., categorical)data types,with missing values all assigned to a special “missing” category”, thus determining missing values is disclosed because Pearson teaches recognizing when data values are missing and further teaches handling incomplete variables by assigning missing values to a special “missing” category. Pearson’s recognition of missing/disguised missing data and identification of incomplete variables correspond to determining missing values) Pearson further teaches determine whether the missing values are informative based on a statistical significance of a category […] (Pearson, Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus determining whether the missing values are informative based on a statistical significance of a category is disclosed because Pearson teaches comparing analysis results when missing/disguised missing values are included versus omitted and determining statistical significance using a t-statistic and p-value. Pearson further teaches that patients with missing serum insulin values differ from patients without missing values, raising the possibility of nonignorable missing data. Pearson’s diabetic/nondiabetic groups and missing/non-missing groups correspond to categories, while the p-value and statistically significant difference correspond to determining whether the missing values are informative based on statistical significance) It would have been obvious to combine Gulin and Guillaume with Pearson because Pearson teaches that disguised missing data can be misinterpreted as valid data and can significantly distort statistical analysis results, while Gulin and Guillaume teach determining split conditions and processing categorical variables for decision tree generation. Pearson further teaches identifying missing values, assigning incomplete variables to categorical missing groups, and using imputation techniques when missing data is informative. Therefore, a POSITA would have been motivated to incorporate Pearson’s missing value detection and handling techniques into the categorical split processing framework of Gulin and Guillaume to improve the reliability and accuracy of categorical split generation and decision tree training when incomplete or disguised missing data is present (Pearson, Page 2 – Section 2, “The problem of disguised missing data arises when missing data values are not explicitly represented as such, but are coded with values that can be misinterpreted as valid data”, & Page 4 – Section 3.2, “If we include the zeros, failing to recognize them as disguised missing data values, we conclude there is no significant difference: the t-statistic has a value of t = 1.80 with 766 degrees of freedom, giving a p-value of 7.2%, not significant at the standard 5% level. Conversely, if we omit the 35 records with zero values for diastolic blood pressure from this analysis, the t-statistic has the value t = 4.68 with 731 degrees of freedom, corresponding to an extremely significant p value of less than 10−16. It is worth emphasizing that this difference is caused by the handling of 5% of the data values, again demonstrating that the presence of even a small concentration of disguised missing data values can have serious consequence”) . Regarding Claim 20, Gulin and Guillaume combined with Pearson teaches all the limitations of claim 19 as cited above and Pearson further teaches: impute the missing values responsive to the missing values being informative (Pearson, Page 2 – Section 1.3, “For larger fractions of missing data, or in other cases where deletion strategies are deemed undesirable, one common alternative is imputation, where missing data values are estimated on the basis of those that are available”, & Page 6 – Section 3.5, “The key point is that the age distribution appears to be different between the two groups: patients with missing serum insulin values appear generally younger than those without missing values. Depending on the analysis undertaken, this difference in patient ages could be important, raising the possibility of nonignorable missing data”, thus imputing the missing values responsive to the missing values being informative is disclosed because Pearson teaches imputing missing data values by estimating them based on available data and further teaches that differences between groups with missing and non-missing values may be important, raising the possibility of nonignorable missing data. Pearson’s imputation of missing data values corresponds to imputing the missing values, while the determination that the missing values may be important/nonignorable corresponds to the missing values being informative) Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US 12299552 is pertinent because it teaches training and adjusting tree-based machine-learning models using decision trees, splitting rules, independent variables, and response variables to generate predicted outputs and explanatory data. The reference further teaches generating partitions and splits within decision trees, modifying splitting rules during model training, enforcing monotonic relationships between variables and predicted responses, and processing categorical and independent-variable data within tree based machine learning frameworks. Because applicant’s disclosure similarly concerns machine learning decision tree models, categorical variable processing, split generation, node assignments, and event based model training and adjustment, the reference is relevant to the invention but is not relied upon in the rejection . Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.A./ Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123 Application/Control Number: 18/501,886 Page 2 Art Unit: 2123 Application/Control Number: 18/501,886 Page 3 Art Unit: 2123 Application/Control Number: 18/501,886 Page 4 Art Unit: 2123 Application/Control Number: 18/501,886 Page 5 Art Unit: 2123 Application/Control Number: 18/501,886 Page 6 Art Unit: 2123 Application/Control Number: 18/501,886 Page 7 Art Unit: 2123 Application/Control Number: 18/501,886 Page 8 Art Unit: 2123 Application/Control Number: 18/501,886 Page 9 Art Unit: 2123 Application/Control Number: 18/501,886 Page 10 Art Unit: 2123 Application/Control Number: 18/501,886 Page 11 Art Unit: 2123 Application/Control Number: 18/501,886 Page 12 Art Unit: 2123 Application/Control Number: 18/501,886 Page 13 Art Unit: 2123 Application/Control Number: 18/501,886 Page 14 Art Unit: 2123 Application/Control Number: 18/501,886 Page 15 Art Unit: 2123 Application/Control Number: 18/501,886 Page 16 Art Unit: 2123 Application/Control Number: 18/501,886 Page 17 Art Unit: 2123 Application/Control Number: 18/501,886 Page 18 Art Unit: 2123 Application/Control Number: 18/501,886 Page 19 Art Unit: 2123 Application/Control Number: 18/501,886 Page 20 Art Unit: 2123 Application/Control Number: 18/501,886 Page 21 Art Unit: 2123 Application/Control Number: 18/501,886 Page 22 Art Unit: 2123 Application/Control Number: 18/501,886 Page 23 Art Unit: 2123 Application/Control Number: 18/501,886 Page 24 Art Unit: 2123 Application/Control Number: 18/501,886 Page 25 Art Unit: 2123 Application/Control Number: 18/501,886 Page 26 Art Unit: 2123 Application/Control Number: 18/501,886 Page 27 Art Unit: 2123 Application/Control Number: 18/501,886 Page 28 Art Unit: 2123 Application/Control Number: 18/501,886 Page 29 Art Unit: 2123 Application/Control Number: 18/501,886 Page 30 Art Unit: 2123 Application/Control Number: 18/501,886 Page 31 Art Unit: 2123 Application/Control Number: 18/501,886 Page 32 Art Unit: 2123 Application/Control Number: 18/501,886 Page 33 Art Unit: 2123 Application/Control Number: 18/501,886 Page 34 Art Unit: 2123 Application/Control Number: 18/501,886 Page 35 Art Unit: 2123 Application/Control Number: 18/501,886 Page 36 Art Unit: 2123 Application/Control Number: 18/501,886 Page 37 Art Unit: 2123 Application/Control Number: 18/501,886 Page 38 Art Unit: 2123 Application/Control Number: 18/501,886 Page 39 Art Unit: 2123 Application/Control Number: 18/501,886 Page 40 Art Unit: 2123 Application/Control Number: 18/501,886 Page 41 Art Unit: 2123 Application/Control Number: 18/501,886 Page 42 Art Unit: 2123 Application/Control Number: 18/501,886 Page 43 Art Unit: 2123 Application/Control Number: 18/501,886 Page 44 Art Unit: 2123 Application/Control Number: 18/501,886 Page 45 Art Unit: 2123 Application/Control Number: 18/501,886 Page 46 Art Unit: 2123 Application/Control Number: 18/501,886 Page 47 Art Unit: 2123 Application/Control Number: 18/501,886 Page 48 Art Unit: 2123 Application/Control Number: 18/501,886 Page 49 Art Unit: 2123 Application/Control Number: 18/501,886 Page 50 Art Unit: 2123 Application/Control Number: 18/501,886 Page 51 Art Unit: 2123 Application/Control Number: 18/501,886 Page 52 Art Unit: 2123 Application/Control Number: 18/501,886 Page 53 Art Unit: 2123 Application/Control Number: 18/501,886 Page 54 Art Unit: 2123 Application/Control Number: 18/501,886 Page 55 Art Unit: 2123 Application/Control Number: 18/501,886 Page 56 Art Unit: 2123 Application/Control Number: 18/501,886 Page 57 Art Unit: 2123 Application/Control Number: 18/501,886 Page 58 Art Unit: 2123 Application/Control Number: 18/501,886 Page 59 Art Unit: 2123 Application/Control Number: 18/501,886 Page 60 Art Unit: 2123 Application/Control Number: 18/501,886 Page 61 Art Unit: 2123 Application/Control Number: 18/501,886 Page 62 Art Unit: 2123 Application/Control Number: 18/501,886 Page 63 Art Unit: 2123 Application/Control Number: 18/501,886 Page 64 Art Unit: 2123 Application/Control Number: 18/501,886 Page 65 Art Unit: 2123 Application/Control Number: 18/501,886 Page 66 Art Unit: 2123 Application/Control Number: 18/501,886 Page 67 Art Unit: 2123 Application/Control Number: 18/501,886 Page 68 Art Unit: 2123
Read full office action

Prosecution Timeline

Nov 03, 2023
Application Filed
May 28, 2026
Non-Final Rejection mailed — §101, §103
Jul 31, 2026
Examiner Interview (Telephonic)
Jul 31, 2026
Examiner Interview Summary

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 5m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month