DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on February 23rd, 2026 has been entered.
Claims 1, 11 and 18 have been amended, claim 21 has been added, and claim 17 has been cancelled. The amendments have been entered, and claims 1-16 and 18-21 are currently pending in the case. Claims 1, 11 and 18 are independent claims.
Claim Rejections - 35 USC § 101
35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 and 18-21 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1: Claim 1 is directed to [a] computer-implemented method, therefore it falls under the statuary category of a method.
Step 2A Prong 1: The claim recites, in part:
“generate a predictive output that describes a likelihood that the predictive input is associated with a target class of a plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III).
“(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“initiating…performance of one or more prediction-based actions based at least in part on the predictive output” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case observation and judgement. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows:
“receiving…a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
“by one or more processors”, “by the one or more processors”, “inputting…the predictive input to the machine learning model to”, “by the one or more processors”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Step 2B: The additional elements “by one or more processors”, “by the one or more processors”, “inputting…the predictive input to the machine learning model to”, “by the one or more processors”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation”, taken individually and in combination, do not provide an inventive concept of significantly more than the abstract idea itself for the reasons set forth in step 2A prong 2 above. Furthermore, “receiving…a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as Furthermore the additional element is directed to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II). Therefore, the claim is ineligible.
Regarding claim 2, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the cumulative target score is determined based at least in part on a combination of a plurality of target scores associated with the set of optimal imbalance adjustment conditions” This encompasses a mathematical concept.
“the cumulative non-target score is determined based at least in part on a combination of a plurality of non-target scores associated with the set of optimal imbalance adjustment conditions” This encompasses the mental determination of a target score based on observed non-target scores.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 3, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the candidate imbalance adjustment condition is associated with an integer non-target score” this encompasses the mental association of candidate imbalance adjustment conditions with integer non-target scores.
“that is determined by mapping the non-target score for the candidate imbalance adjustment condition to a nearest integer” This limitation is a mathematical concept.
“the cumulative non-target score is determined based at least in part on a plurality of integer non-target scores associated with the set of optimal imbalance adjustment condition” this limitation is a mathematical concept.
“maximizing the cumulative target score while the cumulative non-target score satisfies the upper cumulative non-target score threshold comprises using a Knapsack optimization routine” This limitation is a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 4, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“identifying a target subset of the plurality of candidate training entries that are associated with the target class” This encompasses the mental identification of a subset of observed data.
“determining a plurality of per-target-entry condition satisfaction ratios corresponding to the plurality of candidate training entries based at least in part on: (i) a condition satisfaction indicator that describes whether a candidate training entry of the plurality of candidate training entries satisfies the candidate imbalance adjustment condition, and (ii) a cumulative condition satisfaction indicator for the candidate training entry that describes a count of the plurality of candidate imbalance adjustment conditions that are satisfied by the candidate training entry” This limitation can be considered a mathematical concept.
“determining the target score based at least in part on the plurality of per-target-entry condition satisfaction ratio.” This encompasses the mental determination of a target score based on observed data. Further, this limitation can be considered a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 5, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“identifying a non-target subset of the plurality of candidate training entries that are not associated with the target class” this encompasses the mental identification of non-target data amongst observed data.
“determining a plurality of per-non-target-entry condition satisfaction ratios corresponding to the plurality of candidate training entries based at least in part on: (i) a condition satisfaction indicator that describes whether a candidate training entry of the plurality of candidate training entries satisfies the candidate imbalance adjustment condition, and (ii) a cumulative condition satisfaction indicator for the candidate training entry that describes a count of the plurality of candidate imbalance adjustment conditions that are satisfied by the candidate training entry;” This limitation can be considered a mathematical concept.
“determining the target score based at least in part on each per-non-target-entry condition satisfaction ratio” This encompasses the mental determination of a target score based on observed data. Further, this limitation can be considered a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 6, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the target class corresponds to a dependent event that is condition upon occurrence of a primary event” This limitation, under its broadest reasonable interpretation, is a mathematical concept.
“generate a dependent event likelihood for the predictive input with respect to the dependent event and a primary event likelihood for the predictive input with respect to the primary event” this encompasses the mental process of creating a likelihood based on observed events.
“the predictive output is determined based at least in part on the primary event likelihood and the dependent event likelihood” this encompasses the mental determination of output data based on observed event likelihoods. Further, this limitation can be considered a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows:
“the machine learning model is configured to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “the machine learning model is configured to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Regarding claim 7, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“determining an adjusted dependent event likelihood based at least in part on the dependent event likelihood and a dependent event likelihood adjustment parameter” this encompasses the mental determination of an event likelihood based on observed events. Under its broadest reasonable interpretation this could also be considered a mathematical concept.
“determining an adjusted primary event likelihood based at least in part on the primary event likelihood and a primary event likelihood adjustment parameter” this encompasses the mental determination of an event likelihood based on observed events. Under its broadest reasonable interpretation this could also be considered a mathematical concept.
“determining a likelihood product factor based at least in part on the primary event likelihood and the dependent event likelihood” this encompasses the mental determination of an event likelihood based on observed events. Under its broadest reasonable interpretation this could also be considered a mathematical concept.
“determining an adjusted likelihood product factor based at least in part on the likelihood product factor and a likelihood product factor adjustment parameter” this encompasses the mental determination of an event likelihood based on observed events. Under its broadest reasonable interpretation this could also be considered a mathematical concept.
“determining the predictive output based at least in part on the adjusted dependent event likelihood, the adjusted primary event likelihood, and the adjusted likelihood product factor” this encompasses the mental the mental determination of a prediction based on observed events and data.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 8, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter are determined in a manner such that a sum of the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter has a defined summation value” This encompasses a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 9, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter are determined in a manner that is configured to optimize a validation error measure.” This limitation is a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 10, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“the validation error measure is determined based at least in part on a per-entry error measure for each validation entry of one or more validation entries.” This limitation is a mathematical concept.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 11:
Step 1: Claim 11 is directed to [a] system, therefore it falls under the statuary category of a machine.
Step 2A Prong 1: The claim recites, in part:
“generate a predictive output that describes a likelihood that the predictive input is associated with a target class of a plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III).
“(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“initiating performance of one or more prediction-based actions based at least in part on the predictive output” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case observation and judgement. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows:
“receiving a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
“one or more processors”, “when executed by any one or more of the one or more processors, causes the one or more processors to perform operations”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
“at least one memory storing processor-executable instructions” the limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
Step 2B: The additional elements “one or more processors”, “when executed by any one or more of the one or more processors, causes the one or more processors to perform operations”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation”, “at least one memory storing processor-executable instructions”, taken individually and in combination, do not provide an inventive concept of significantly more than the abstract idea itself for the reasons set forth in step 2A prong 2 above. Furthermore, “receiving a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as Furthermore the additional element is directed to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II). Therefore, the claim is ineligible.
Regarding claims 12-16:
The rejection of claim 11 is further incorporated, the rejection of claims 2-6 are applicable to claims 12-16, respectively.
Regarding claim 18:
Step 1: Claim 18 is directed to [o]ne or more non-transitory computer-readable storage media, therefore it falls under the statuary category of a manufacture.
Step 2A Prong 1: The claim recites, in part:
“generate a predictive output that describes a likelihood that the predictive input is associated with a target class of a plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III).
“(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III). Further, this limitation is a mathematical concept. See MPEP § 2106.04(a)(2)( I).
“initiating performance of one or more prediction-based actions based at least in part on the predictive output” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case observation and judgement. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows:
“receiving a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
“one or more processors”, “when executed by any one or more of the one or more processors, causes the one or more processors to”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Step 2B: The additional elements “one or more processors”, “when executed by any one or more of the one or more processors, causes the one or more processors to”, “(iii) training the machine learning model by updating one or more parameters of the machine learning model based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation” taken individually and in combination, do not provide an inventive concept of significantly more than the abstract idea itself for the reasons set forth in step 2A prong 2 above. Furthermore, “receiving a predictive input for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as Furthermore the additional element is directed to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II). Therefore, the claim is ineligible.
Regarding claims 19-20:
The rejection of claim 18 is further incorporated, the rejection of claims 2-3 are applicable to claims 19-20, respectively.
Regarding claim 21, the rejection of the parent claim is incorporated and further:
Step 1: A process, as identified in independent claim 1.
Step 2A Prong 1: The claim recites, in part:
“(a) generating the first training entry batch, wherein a training entry of the first training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the primary event” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III).
“(c) generating a second training entry batch, wherein a training entry of the second training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the dependent event” This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case evaluation. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows:
“(i) the machine learning model comprises an identifier machine learning node and a differentiator machine learning node”, “(ii) the machine learning model is trained using a defined number of training iterations” these limitations are an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
“(b) updating parameters of the identifier machine learning node using the first training entry batch”, “(d) updating parameters of the differentiator machine learning node using the second training entry batch” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “(i) the machine learning model comprises an identifier machine learning node and a differentiator machine learning node”, “(ii) the machine learning model is trained using a defined number of training iterations” these limitations are an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
“(b) updating parameters of the identifier machine learning node using the first training entry batch”, “(d) updating parameters of the differentiator machine learning node using the second training entry batch” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 5, 11, 12, 14, 15, 18 and 19 are rejected under 35 U.S.C. § 103 as being unpatentable over Kumar ("Extending Active Learning for Improved Long-Term Return On Investment of Learning Systems", Kumar, 2013) in view of Dong et al. ("CAEP: Classification by Aggregating Emerging Patterns", Dong et al., 01 December 1999) (hereinafter "Dong") in further view of Zhao et al. (“Self-interpretable Convolutional Neural Networks for Text Classification”, Zhao et al., May 14, 2021) (hereinafter “Zhao)”
Regarding claim 1:
Kumar teaches computer-implemented method comprising:
receiving, by one or more processors (Kumar, page 84, ¶3 “The system has been implemented in Matlab® and the experiments are run on two 32-processor Intel® Xeon® E5-2450 @ 2.10GHz machines with 1 Tb RAM running Ubuntu OS in the Carnegie Mellon University Speech cluster.”), a predictive input (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance.” Here the health insurance claims can be considered the predictive input)
inputting, by the one or more processors (Kumar, page 84, ¶3 “The system has been implemented in Matlab® and the experiments are run on two 32-processor Intel® Xeon® E5-2450 @ 2.10GHz machines with 1 Tb RAM running Ubuntu OS in the Carnegie Mellon University Speech cluster.”), the predictive input to the machine learning model (Kumar, page 5, ¶1 “In the last chapter 7, we mention the details of our work on cost-sensitive exploitation and deploying the machine learning system for Claims Rework domain.”) to generate a predictive output that describes a likelihood that the predictive input is associated with a target class (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance. The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.” Here, the likelihood a claim has an errors can be considered a likelihood that the predictive input is associated with a target class) of a plurality of candidate classes (Kumar, page 42, ¶3 “We create data samples of 1100 claims per iteration with 5% error(55) and 95% correct(1045) claims.” , “error” and “correct” can be considered the plurality of candidate classes), wherein the machine learning model is trained by (Kumar, page 25, ¶2 “Thus equation (3.5) gives the ‘expected’ utility for adding the unlabeled example x to the training data”):
(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if we consider that the cost of the set is a simple aggregation of the element-wise cost then the costs of the different sets are S1:4, S2:10, S3:6, S4:10, S5:10. If the budget given is 14 then the optimal selections would be S1 and S5.” Here, S1 and S5 are selected from the plurality of sets. Each set contains candidate imbalance conditions, Kumar, page 21, ¶3 “each example can potentially have a different value that is a function of the example’s (or population’s) features”) in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions Kumar, page 27, ¶1 “In reference to figure 3.1, if k=2 then the optimal set coverage is with S3 and S4 which covers all the elements. In the weighted version every element
e
j
has a weight
w
(
e
j
)
. The task is to find a maximum coverage which has maximum weight.” Here the value of the weights can be considered the target score) while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold (Kumar, page 27, ¶1 “In the budgeted maximum coverage version, not only does every element
e
j
has a weight
w
(
e
j
)
, but also every set
S
i
has a cost
c
(
S
i
)
. Instead of k that limits the number of sets, in the budgeted version a budget B is given. This budget B limits the weight of the cover that can be chosen.” Here, the total cost can be considered the non-target score and the budget can be considered the non-target threshold), wherein:
(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes (Kumar, page 88, ¶2 “We sample equally from most confident positive as well as negative examples in order to come up with a balanced training dataset.”), and
(iii) training the machine learning model by updating one or more parameters of the machine learning model (Kumar, page 79, section 5.4.5, ¶2 “We have used the liblinear package [Fan et al., 2008] for classification and use logistic regression as the learning algorithm.”) based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation (Kumar, page 79, section 5.4.5, ¶1 “Thus the only strategy to deal with class imbalance that is relevant in the context of the thesis and the interactive framework is to modify the misclassification costs per class.” Here, the dealing with class imbalance can be considered the removing of class imbalance performance degradation in light of “There are three major approaches in machine learning literature [Japkowicz and Stephen, 2002] that are used for dealing with class imbalance and improving classification performance, namely, oversampling minority class, undersampling majority class and modifying misclassification costs per class.”); and
initiating, by the one or more processors, performance of one or more prediction-based actions based at least in part on the predictive output (Kumar, page 42, ¶2 “The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.”)
Kumar does not teach "(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes"
However, Dong teaches (a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2% ).” Here the items support (frequency) for edible can be considered the target score comprising a correlation measure between an item and a target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”)
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2%).” Here the items support (frequency) for poisonous can be considered the non-target score comprising a correlation measure between an item and a non-target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”) and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes (Dong, “To illustrate this point, consider these two EPs of the Iris-versicolor class from the Iris dataset [12]: e1 = ({1,5,11},3%,∞) e2 = ({11},100%,22.25) e2 is clearly more useful than e1 for classification, since it covers 32 times more instances and its associated odds, 95.7%, is also very near that of the other EP, e1.”)
Kumar and Dong are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar’s deep learning prediction system to incorporate the target and non-target scores taught by Dong. The motivation for doing so would have been to perform efficiently and accurately on all classes and datasets, as stated in Zhao, page 2, ¶3 “The resulting classier CAEP is in general more accurate than C4.5 and CBA, and is a lot more accurate than them over datasets where they do not have good performance. CAEP is equally accurate on all classes, and can be built efficiently from large, even high dimensional datasets.”
Kumar in view of Dong does not teach " for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers"
However, Zhao teaches for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers (Zhao, page 13, section 4.2 “The interesting feature engineering property of convolutional layer and local linearity property of ReLU DNN classifier combined together provides us easy and intuitive interpretation of the CNN model” here, the feature engineering property of convolutional layer can be considered the one or more feature engineering layers and the ReLU DNN classifier can be considered one or more feature processing layers);
Kumar in view of Dong and Zhao are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar/Dong’s deep learning prediction system to incorporate the feature processing and engineering layers taught by Zhao. The motivation for doing so would have been to achieve very good predictive performances, as stated in Zhao, page 2, ¶3 “The convolutional layer does feature engineering, which selects important information from the text documents, and is interpretable when used in conjunction with max-pooling layer. We apply regularization on the ReLU DNN after the convolutional layer to reduce model complexity. Our experiments demonstrate such regularized model can achieve very good predictive performances and are simpler for model interpretation.”
Regarding claim 2:
Kumar in view of Dong in further view of Zhao teaches [t]he computer-implemented method of claim 1, wherein:
the cumulative target score is determined based at least in part on a combination of a plurality of target scores associated with the set of optimal imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if k=2 then the optimal set coverage is with S3 and S4 which covers all the elements. In the weighted version every element
e
j
has a weight
w
(
e
j
)
. The task is to find a maximum coverage which has maximum weight.” Here the value of the weights can be considered the target score and the sum of the weights which is the maximum weight, is the cumulative target score), and
the cumulative non-target score is determined based at least in part on a combination of a plurality of non-target scores associated with the set of optimal imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if we consider that the cost of the set is a simple aggregation of the element-wise cost then the costs of the different sets are S1:4, S2:10, S3:6, S4:10, S5:10. If the budget given is 14 then the optimal selections would be S1 and S5” the aggregation of the costs can be considered a cumulative non-target score).
Regarding claim 4:
Kumar in view of Dong in further view of Zhao teaches [t]he computer-implemented method of claim 2, wherein determining the target score for the candidate imbalance adjustment condition comprises:
identifying a target subset of the plurality of candidate training entries that are associated with the target class (Kumar, page 29, algorithm 1
PNG
media_image1.png
343
590
media_image1.png
Greyscale
“(a) U ← S\G” here, U can be considered a subset of the entries of the target class);
determining a plurality of per-target-entry condition satisfaction ratios corresponding to the plurality of candidate training entries (Kumar, page 29, algorithm 1 “i. select
s
i
∈
U
that maximizes
W
i
'
c
i
”) based at least in part on:
(i) a condition satisfaction indicator that describes whether a candidate training entry of the plurality of candidate training entries satisfies the candidate imbalance adjustment condition (Kumar, page 28, ¶2 “A collection of sets S = {S1, S2,...Sm}”), and
(ii) a cumulative condition satisfaction indicator for the candidate training entry that describes a count of the plurality of candidate imbalance adjustment conditions that are satisfied by the candidate training entry (Kumar, page 29, ¶2 “Let
W
i
'
,
ⅈ
=
1
,
…
m
, denote the total weight of the elements covered by set Si , but not covered by any set in G.”); and
determining the target score based at least in part on [[each]] the plurality of per-target- entry condition satisfaction ratios (Kumar, page 29, algorithm 1 “4. if w(H1) > w(H2), output H1, otherwise, output H2” as this is the weight of H1/H2 it can be considered the target score, and determined, in part, by the ratio).
Regarding claim 5:
Kumar in view of Dong in further view of Zhao teaches [t]he computer-implemented method of claim 2, wherein determining the non-target score for the candidate imbalance adjustment condition comprises:
identifying a non-target subset of the plurality of candidate training entries that are not associated with the target class (Kumar, page 29, algorithm 1
PNG
media_image1.png
343
590
media_image1.png
Greyscale
“(a) U ← S\G” here, U can be considered a subset of the entries);
determining a plurality of per-non-target-entry condition satisfaction ratios corresponding to the plurality of candidate training entries (Kumar, page 29, algorithm 1 “i. select
s
i
∈
U
that maximizes
W
i
'
c
i
”) based at least in part on:
(i) a condition satisfaction indicator that describes whether a candidate training entry of the plurality of candidate training entries satisfies the candidate imbalance adjustment condition (Kumar, page 28, ¶2 “A collection of sets S = {S1, S2,...Sm} with associated costs
m
i
ⅈ
=
1
m
”), and
(ii) a cumulative condition satisfaction indicator for the candidate training entry that describes a count of the plurality of candidate imbalance adjustment conditions that are satisfied by the candidate training entry (Kumar, page 29, ¶2 “Let
W
i
'
,
ⅈ
=
1
,
…
m
, denote the total weight of the elements covered by set Si , but not covered by any set in G.”); and
determining the target score based at least in part on the plurality of per-non-target-entry condition satisfaction ratios (Kumar, page 29, algorithm 1 “4. if w(H1) > w(H2), output H1, otherwise, output H2” as this is the weight of H1/H2 it can be considered the target score, and determined, in part, by the ratio).
Regarding claim 11:
Kumar teaches A system comprising:
at least one memory storing processor-executable instructions that, when executed by any one or more of the one or more processors, causes the one or more processors to perform operations (Kumar, page 84, ¶3 “The system has been implemented in Matlab® and the experiments are run on two 32-processor Intel® Xeon® E5-2450 @ 2.10GHz machines with 1 Tb RAM running Ubuntu OS in the Carnegie Mellon University Speech cluster.”) comprising:
receive a predictive input (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance.” Here the health insurance claims can be considered the predictive input)
input, the predictive input to the machine learning model (Kumar, page 5, ¶1 “In the last chapter 7, we mention the details of our work on cost-sensitive exploitation and deploying the machine learning system for Claims Rework domain.”) to generate a predictive output that describes a likelihood that the predictive input is associated with a target class (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance. The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.” Here, the likelihood a claim has an errors can be considered a likelihood that the predictive input is associated with a target class) of a plurality of candidate classes (Kumar, page 42, ¶3 “We create data samples of 1100 claims per iteration with 5% error(55) and 95% correct(1045) claims.” , “error” and “correct” can be considered the plurality of candidate classes), wherein the machine learning model is trained by (Kumar, page 25, ¶2 “Thus equation (3.5) gives the ‘expected’ utility for adding the unlabeled example x to the training data”):
(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if we consider that the cost of the set is a simple aggregation of the element-wise cost then the costs of the different sets are S1:4, S2:10, S3:6, S4:10, S5:10. If the budget given is 14 then the optimal selections would be S1 and S5.” Here, S1 and S5 are selected from the plurality of sets. Each set contains candidate imbalance conditions, Kumar, page 21, ¶3 “each example can potentially have a different value that is a function of the example’s (or population’s) features”) in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions Kumar, page 27, ¶1 “In reference to figure 3.1, if k=2 then the optimal set coverage is with S3 and S4 which covers all the elements. In the weighted version every element
e
j
has a weight
w
(
e
j
)
. The task is to find a maximum coverage which has maximum weight.” Here the value of the weights can be considered the target score) while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold (Kumar, page 27, ¶1 “In the budgeted maximum coverage version, not only does every element
e
j
has a weight
w
(
e
j
)
, but also every set
S
i
has a cost
c
(
S
i
)
. Instead of k that limits the number of sets, in the budgeted version a budget B is given. This budget B limits the weight of the cover that can be chosen.” Here, the total cost can be considered the non-target score and the budget can be considered the non-target threshold), wherein:
(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes (Kumar, page 88, ¶2 “We sample equally from most confident positive as well as negative examples in order to come up with a balanced training dataset.”), and
(iii) training the machine learning model by updating one or more parameters of the machine learning model (Kumar, page 79, section 5.4.5, ¶2 “We have used the liblinear package [Fan et al., 2008] for classification and use logistic regression as the learning algorithm.”) based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation (Kumar, page 79, section 5.4.5, ¶1 “Thus the only strategy to deal with class imbalance that is relevant in the context of the thesis and the interactive framework is to modify the misclassification costs per class.” Here, the dealing with class imbalance can be considered the removing of class imbalance performance degradation in light of “There are three major approaches in machine learning literature [Japkowicz and Stephen, 2002] that are used for dealing with class imbalance and improving classification performance, namely, oversampling minority class, undersampling majority class and modifying misclassification costs per class.”); and
Initiate performance of one or more prediction-based actions based at least in part on the predictive output (Kumar, page 42, ¶2 “The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.”)
Kumar does not teach "(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes"
However, Dong teaches (a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2% ).” Here the items support (frequency) for edible can be considered the target score comprising a correlation measure between an item and a target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”)
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2%).” Here the items support (frequency) for poisonous can be considered the non-target score comprising a correlation measure between an item and a non-target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”) and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes (Dong, “To illustrate this point, consider these two EPs of the Iris-versicolor class from the Iris dataset [12]: e1 = ({1,5,11},3%,∞) e2 = ({11},100%,22.25) e2 is clearly more useful than e1 for classification, since it covers 32 times more instances and its associated odds, 95.7%, is also very near that of the other EP, e1.”)
Kumar and Dong are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar’s deep learning prediction system to incorporate the target and non-target scores taught by Dong. The motivation for doing so would have been to perform efficiently and accurately on all classes and datasets, as stated in Zhao, page 2, ¶3 “The resulting classier CAEP is in general more accurate than C4.5 and CBA, and is a lot more accurate than them over datasets where they do not have good performance. CAEP is equally accurate on all classes, and can be built efficiently from large, even high dimensional datasets.”
Kumar in view of Dong does not teach " for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers"
However, Zhao teaches for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers (Zhao, page 13, section 4.2 “The interesting feature engineering property of convolutional layer and local linearity property of ReLU DNN classifier combined together provides us easy and intuitive interpretation of the CNN model” here, the feature engineering property of convolutional layer can be considered the one or more feature engineering layers and the ReLU DNN classifier can be considered one or more feature processing layers);
Kumar in view of Dong and Zhao are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar/Dong’s deep learning prediction system to incorporate the feature processing and engineering layers taught by Zhao. The motivation for doing so would have been to achieve very good predictive performances, as stated in Zhao, page 2, ¶3 “The convolutional layer does feature engineering, which selects important information from the text documents, and is interpretable when used in conjunction with max-pooling layer. We apply regularization on the ReLU DNN after the convolutional layer to reduce model complexity. Our experiments demonstrate such regularized model can achieve very good predictive performances and are simpler for model interpretation.”
Regarding claims 12, 14 and 15 they are rejected under the same rational as claims 2, 4 and 5 respectively.
Regarding claim 18:
Kumar teaches [o]ne or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors (Kumar, page 84, ¶3 “The system has been implemented in Matlab® and the experiments are run on two 32-processor Intel® Xeon® E5-2450 @ 2.10GHz machines with 1 Tb RAM running Ubuntu OS in the Carnegie Mellon University Speech cluster.”), cause the one or more processors to:
receive a predictive input (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance.” Here the health insurance claims can be considered the predictive input)
input, the predictive input to the machine learning model (Kumar, page 5, ¶1 “In the last chapter 7, we mention the details of our work on cost-sensitive exploitation and deploying the machine learning system for Claims Rework domain.”) to generate a predictive output that describes a likelihood that the predictive input is associated with a target class (Kumar, page 42, ¶2 “This task deals with the problem of predicting (and reducing) payment errors when processing health insurance claims which belongs to the same class of problems as fraud detection, intrusion detection, and surveillance. The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.” Here, the likelihood a claim has an errors can be considered a likelihood that the predictive input is associated with a target class) of a plurality of candidate classes (Kumar, page 42, ¶3 “We create data samples of 1100 claims per iteration with 5% error(55) and 95% correct(1045) claims.” , “error” and “correct” can be considered the plurality of candidate classes), wherein the machine learning model is trained by (Kumar, page 25, ¶2 “Thus equation (3.5) gives the ‘expected’ utility for adding the unlabeled example x to the training data”):
(i) determining a set of optimal imbalance adjustment conditions from a plurality of candidate imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if we consider that the cost of the set is a simple aggregation of the element-wise cost then the costs of the different sets are S1:4, S2:10, S3:6, S4:10, S5:10. If the budget given is 14 then the optimal selections would be S1 and S5.” Here, S1 and S5 are selected from the plurality of sets. Each set contains candidate imbalance conditions, Kumar, page 21, ¶3 “each example can potentially have a different value that is a function of the example’s (or population’s) features”) in a manner that is configured to maximize a cumulative target score for the set of optimal imbalance adjustment conditions Kumar, page 27, ¶1 “In reference to figure 3.1, if k=2 then the optimal set coverage is with S3 and S4 which covers all the elements. In the weighted version every element
e
j
has a weight
w
(
e
j
)
. The task is to find a maximum coverage which has maximum weight.” Here the value of the weights can be considered the target score) while a cumulative non-target score for the set of optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold (Kumar, page 27, ¶1 “In the budgeted maximum coverage version, not only does every element
e
j
has a weight
w
(
e
j
)
, but also every set
S
i
has a cost
c
(
S
i
)
. Instead of k that limits the number of sets, in the budgeted version a budget B is given. This budget B limits the weight of the cover that can be chosen.” Here, the total cost can be considered the non-target score and the budget can be considered the non-target threshold), wherein:
(ii) determining a set of filtered training entries from the plurality of candidate training entries, wherein the set of filtered training entries comprise a balanced ratio of entries associated with the target class compared to the other classes of the plurality of candidate classes (Kumar, page 88, ¶2 “We sample equally from most confident positive as well as negative examples in order to come up with a balanced training dataset.”), and
(iii) training the machine learning model by updating one or more parameters of the machine learning model (Kumar, page 79, section 5.4.5, ¶2 “We have used the liblinear package [Fan et al., 2008] for classification and use logistic regression as the learning algorithm.”) based at least in part on a first training entry batch that is filtered using the set of filtered training entries to remove class imbalance performance degradation (Kumar, page 79, section 5.4.5, ¶1 “Thus the only strategy to deal with class imbalance that is relevant in the context of the thesis and the interactive framework is to modify the misclassification costs per class.” Here, the dealing with class imbalance can be considered the removing of class imbalance performance degradation in light of “There are three major approaches in machine learning literature [Japkowicz and Stephen, 2002] that are used for dealing with class imbalance and improving classification performance, namely, oversampling minority class, undersampling majority class and modifying misclassification costs per class.”); and
Initiate performance of one or more prediction-based actions based at least in part on the predictive output (Kumar, page 42, ¶2 “The goal is to minimize errors by predicting which claims are likely to have errors, and presenting the highly scored ones to human auditors so that they can be corrected before being finalized and paid.”)
Kumar does not teach "(a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes"
However, Dong teaches (a) a target score for a candidate imbalance adjustment condition of the plurality of candidate imbalance adjustment conditions comprises a first estimated correlation measure between the candidate imbalance adjustment condition and a target subset of a plurality of candidate training entries is associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2% ).” Here the items support (frequency) for edible can be considered the target score comprising a correlation measure between an item and a target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”)
(b) a non-target score for the candidate imbalance adjustment condition comprises a second estimated correlation measure between the candidate imbalance adjustment condition and a non-target subset of the plurality of candidate training entries that is not associated with the target class (Dong, page 2, ¶1 “Roughly speaking, EPs are those item sets whose supports (i.e. frequencies) increase signicantly from one class of data to another. For example, the itemset {odor=none, stalk-surface-below-ring = smooth, ring-number=one} in the Mushroom dataset [12] is a typical EP, whose support increases from 0.2% in the poisonous class to 57.6% in the edible class, at a growth rate of 288 (
57.6
%
0.2
%
= 57.6% 0.2%).” Here the items support (frequency) for poisonous can be considered the non-target score comprising a correlation measure between an item and a non-target subset. Further, Dong, page 5, section 3.1, ¶1 “For each class Ck, we will use a set of EPs to contrast its instances, Dk, against all other instances: We let
D
k
'
=
D
-
D
k
be the opposing class, or simply opponent, of Dk. We then mine (discussion on how is given later) the EPs from
D
k
'
to
D
k
; we refer to these EPs as the EPs of class Ck, and sometimes refer to Ck as the target class of these EPs.”) and comprising an imbalanced ratio of entries associated with the target class compared to other classes of the plurality of candidate classes (Dong, “To illustrate this point, consider these two EPs of the Iris-versicolor class from the Iris dataset [12]: e1 = ({1,5,11},3%,∞) e2 = ({11},100%,22.25) e2 is clearly more useful than e1 for classification, since it covers 32 times more instances and its associated odds, 95.7%, is also very near that of the other EP, e1.”)
Kumar and Dong are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar’s deep learning prediction system to incorporate the target and non-target scores taught by Dong. The motivation for doing so would have been to perform efficiently and accurately on all classes and datasets, as stated in Zhao, page 2, ¶3 “The resulting classier CAEP is in general more accurate than C4.5 and CBA, and is a lot more accurate than them over datasets where they do not have good performance. CAEP is equally accurate on all classes, and can be built efficiently from large, even high dimensional datasets.”
Kumar in view of Dong does not teach " for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers"
However, Zhao teaches for a machine learning model comprising one or more feature engineering layers and one or more feature processing layers (Zhao, page 13, section 4.2 “The interesting feature engineering property of convolutional layer and local linearity property of ReLU DNN classifier combined together provides us easy and intuitive interpretation of the CNN model” here, the feature engineering property of convolutional layer can be considered the one or more feature engineering layers and the ReLU DNN classifier can be considered one or more feature processing layers);
Kumar in view of Dong and Zhao are analogous art because both references concern methods for deep learning predictions. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar/Dong’s deep learning prediction system to incorporate the feature processing and engineering layers taught by Zhao. The motivation for doing so would have been to achieve very good predictive performances, as stated in Zhao, page 2, ¶3 “The convolutional layer does feature engineering, which selects important information from the text documents, and is interpretable when used in conjunction with max-pooling layer. We apply regularization on the ReLU DNN after the convolutional layer to reduce model complexity. Our experiments demonstrate such regularized model can achieve very good predictive performances and are simpler for model interpretation.”
Regarding claim 19, it is rejected under the same rational as claim 2.
Claims 3, 13 and 20 are rejected under 35 U.S.C. § 103 as being unpatentable over Kumar in view of Dong in view of Zhao in further view of Chekuri ("CS 598CSC: Approximation Algorithms", Chekuri, 2009).
Regarding claim 3:
Kumar in view of Dong in further view of Zhao teaches [t]he computer-implemented method of claim 2, wherein:
the cumulative non-target score is determined based at least in part on a plurality of integer non-target scores associated with the set of optimal imbalance adjustment conditions (Kumar, page 27, ¶1 “In reference to figure 3.1, if we consider that the cost of the set is a simple aggregation of the element-wise cost then the costs of the different sets are S1:4, S2:10, S3:6, S4:10, S5:10. If the budget given is 14 then the optimal selections would be S1 and S5”)
Kumar in view of Dong in further view of Zhao does not teach “the candidate imbalance adjustment condition is associated with an integer non-target score that is determined by mapping the non-target score for the candidate imbalance adjustment condition to a nearest integer,
maximizing the cumulative target score while the cumulative non-target score satisfies the upper cumulative non-target score threshold comprises using a Knapsack optimization routine.”
However, Chekuri teaches the candidate imbalance adjustment condition is associated with an integer non- target score that is determined by mapping the non-target score for the candidate imbalance adjustment condition to a nearest integer (Chekuri, page 3, observation 7 “Now, fix some ∈ ∈ (0, 1). We want to scale the profits and round them to be integers so we may use the O(nP) algorithm efficiently while still keeping enough information in the numbers to allow for an accurate approximation.”),
maximizing the cumulative target score while the cumulative non-target score satisfies the upper cumulative non-target score threshold comprises using a Knapsack optimization routine (Chekuri, page 4, Proof “
PNG
media_image2.png
45
468
media_image2.png
Greyscale
The algorithm returns the best choice for A given the scaled and rounded values, so we know
p
'
A
≥
p
'
A
*
It should be noted that this is not the best FPTAS known for Knapsack” here it is shown that the mapped integer values can be used to solve the knapsack optimization).
Kumar in view of Dong in further view of Zhao and Chekuri are analogous art because both references concern methods for Budgeted Max Cover and the knapsack optimization. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar’s knapsack optimization (Kumar, page 28, ¶1 “The non-markovian setups get reduced to the Budgeted Max Cover1 problem easily and are efficiently solved as well” and Kumar, page 28, footnote 1 “Budgeted max coverage reduces to 0/1 knapsack problem with singleton sets and no markovian effect” here it can be seen that the maximization can be considered a knapsack optimization) to incorporate rounding before knapsack optimization as taught by Chekuri. The motivation for doing so would have been to “use the O(nP) algorithm efficiently while still keeping enough information in the numbers to allow for an accurate approximation” Chekuri, page 3, observation 7.
Regarding claims 13 and 20, they are rejected under the same rational as claim 3.
Claims 6-10 and 16 are rejected under 35 U.S.C. § 103 as being unpatentable over Kumar in view of Dong in view of Zhao in view further of Colucci (U.S. Patent US 11,010,848 B1).
Regarding claim 6:
Kumar in view of Dong in further view of Zhao teaches [t]he computer-implemented method of claim 1, wherein:
Kumar in view of Dong in further view of Zhao does not teach “the target class corresponds to a dependent event that is condition upon occurrence of a primary event, the machine learning model is configured to generate a dependent event likelihood for the predictive input with respect to the dependent event and a primary event likelihood for the predictive input with respect to the primary event, and
the predictive output is determined based at least in part on the primary event likelihood and the dependent event likelihood.”
However, Colucci teaches the target class corresponds to a dependent event that is condition upon occurrence of a primary event, the machine learning model is configured to generate a dependent event likelihood for the predictive input with respect to the dependent event (Colucci, col 5, lines 34-35 “In one embodiment, sensitivity tables can be graphically generated and presented to reflect the varying predicted outcomes” here, a sensitivity table can be considered a measure, and the outcome can be considered the dependent event) and a primary event likelihood for the predictive input with respect to the primary event (Colucci, col 4, lines 29-33 “An outcome associated with a predictive model can include a suggested settlement offer, the predicted outcome if the matter were to be tried in a court, a likelihood measure of the settlement offer and/or the predicted outcome.” Here the likelihood measure of the settlement can be considered the primary event), and
the predictive output is determined based at least in part on the primary event likelihood and the dependent event likelihood (Colucci, claim 1 “determining a predicted outcome using the trained predictive model, the matter data and external matter data;” and Colucci, claim 6 “The method for predicting matter outcome of claim 2, wherein the predicted outcome includes a predicted cost of litigation or predicted award.” Here the outcome can be considered the primary event, and the cost of litigation can be considered the dependent event).
Kumar in view of Dong in further view of Zhao and Colucci are analogous art because both references concern methods for predicting outcomes and issues in insurance and legal outcomes. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar/Dong/Zhao’s method to incorporate primary and dependent events as taught by Colucci. The motivation for doing so would have been to have use the model of Kumar to determine the ramifications of combining causes of action “In at least one embodiment the predictive models can be associated with other predictive models and can include an output of the ramifications of combining causes of action. In addition to output for the ramifications of combining the models, output can also be generated from the impact one cause of action will have on the other.” Colucci, col 10, lines 47-52.
Regarding claim 7:
Kumar in view of Dong in view of Zhao in view further of Colucci teaches [t]he computer-implemented method of claim 6, wherein generating the predictive output comprises:
determining an adjusted dependent event likelihood based at least in part on the dependent event likelihood and a dependent event likelihood adjustment parameter (Colucci, col 8, lines 2-7 “In at least one embodiment the relevancy weight of the relevancy factors can be dynamically adjusted based on the information collected. The information collected can include case law, outcomes of previously predicted matters, news reports, legislation or any other relevant information.” Further, Colucci, col 9, lines 38-42 “Relevant trends 404 can also be identified relevant trends and can include case law specific to characteristics of the matter. The relevant trend information can be used to adjust the weights of the factors considered in the predictive model 407” further, Colucci, col 9, lines 45-46 “Relevant history 405 can also be considered by the predictive model to determine an output 408.”),
determining an adjusted primary event likelihood based at least in part on the primary event likelihood and a primary event likelihood adjustment parameter (Colucci, col 8, lines 2-7 “In at least one embodiment the relevancy weight of the relevancy factors can be dynamically adjusted based on the information collected. The information collected can include case law, outcomes of previously predicted matters, news reports, legislation or any other relevant information.” Further, Colucci, col 9, lines 38-42 “Relevant trends 404 can also be identified relevant trends and can include case law specific to characteristics of the matter. The relevant trend information can be used to adjust the weights of the factors considered in the predictive model 407” further, Colucci, col 9, lines 45-46 “Relevant history 405 can also be considered by the predictive model to determine an output 408.”),
determining a likelihood product factor based at least in part on the primary event likelihood and the dependent event likelihood (Colucci, col 9, lines 45-51 “Relevant history 405 can also be considered by the predictive model to determine an output 408. Relevant history can include the information of settlements, damage awards, information about the defendant and/or potential plaintiff's legal history (i.e., settlements, damages, current litigations, etc.). Any other relevant data 406 can also be used by the predictive model to determine an output.”),
determining an adjusted likelihood product factor based at least in part on the likelihood product factor and a likelihood product factor adjustment parameter (Colucci, col 11, lines 20-26 “The training can include adding or removing relevant factors, and/or adjusting the weights of the relevant factors” here it can be seen that the various likelihood factors can be adjusted), and
determining the predictive output based at least in part on the adjusted dependent event likelihood, the adjusted primary event likelihood, and the adjusted likelihood product factor (Colucci, Claim 7 “outputting the settlement offer” dependent on Colucci, claim 2, “updating the predictive model includes adjusting the associated weight of one or more of the relevant factors”).
Regarding claim 8:
Kumar in view of Dong in view of Zhao in view further of Colucci teaches [t]he computer-implemented method of claim 7, wherein the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter are determined in a manner such that a sum of the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter has a defined summation value (Colucci, col 11, lines 20-26 “The training can include adding or removing relevant factors, and/or adjusting the weights of the relevant factors” since the weights of the factors can be adjusted, they can be configured such that they are any defined value.).
Regarding claim 9:
Kumar in view of Dong in view of Zhao in view further of Colucci teaches [t]he computer-implemented method of claim 7, wherein the dependent event likelihood adjustment parameter, the primary event likelihood adjustment parameter, and the likelihood product factor adjustment parameter are determined in a manner that is configured to optimize a validation error measure (Kumar, page 25, ¶3 “In words, equation (3.6) states that the utility function is the increase in the performance for the metric Precision@K that is achieved by adding the example x to the training data. There are multiple ways in which the Precision@K(D) can be evaluated based on the expected error reduction literature for active learning [Baram et al., 2004, Lindenbaum et al., 2004, Roy and Mccallum, 2001].” Here the increase in the performance metric can be considered an optimization).
Regarding claim 10:
Kumar in view of Dong in view of Zhao in further view of Colucci teaches [t]he computer-implemented method of claim 9, wherein the validation error measure is determined based at least in part on a per-entry error measure for each validation entry of one or more validation entries (Kumar, page 25, ¶4 “Labeled training history: Measure the performance of the classifier learnt on the labeled data D so far. This method measures training data error (self-error) and has been pointed out by Baram et al. [2004], Schohn and Cohn [2000] to give substantially biased estimates of the classifiers performance due to the biased sample acquired by the active learner “).
Regarding claim 16, it is rejected under the same rational as claim 6.
Claims 21 is rejected under 35 U.S.C. § 103 as being unpatentable over Kumar in view of Dong in view of Zhao in view further of Colucci in further view of Sarwar et al. ("Two-stage Cascaded Classifier for Purchase Prediction", Sarwar et al., 16 Aug 2015) (hereinafter "Sarwar").
Regarding claim 21:
Kumar in view of Dong in view of Zhao in view further of Colucci teaches [t]he computer-implemented method of claim 6
Kumar in view of Dong in view of Zhao in view further of Colucci does not teach "(i) the machine learning model comprises an identifier machine learning node and a differentiator machine learning node,
(ii) the machine learning model is trained using a defined number of training iterations, and (iii) a training iteration of the defined number of training iterations comprises:
(a) generating the first training entry batch, wherein a training entry of the first training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the primary event,
(b) updating parameters of the identifier machine learning node using the first training entry batch,
(c) generating a second training entry batch, wherein a training entry of the second training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the dependent event, and
(d) updating parameters of the differentiator machine learning node using the second training entry batch"
However, Sarwar teaches (i) the machine learning model comprises an identifier machine learning node (Sarwar, pages 3-4, col 2-1, section 4.2, ¶1 “There are two important observations for building the session classifier: 1) Since we have only 0.05 penalty for selecting a non-buy session, it should be a high recall classifier; 2) As there exists class imbalance problem, a classifier would tend to predict most of the sessions as non-buy.” Here, the session classifier can be considered the identifier node) and a differentiator machine learning node (Sarwar, page 3, col 1, section 4.1, ¶1 “An important observation about item classifier is that we have to train the classifier only with the click data of the buy sessions in training data.” Here, the item classifier can be considered the differentiator node),
(ii) the machine learning model is trained using a defined number of training iterations (Sarwar, page 4, col 1, ¶2 “In order to address the problems we performed resampling of data and classified sessions using AdaBoost.M1 algorithm… For Adaboost.M1 the model is built in a smaller period of time (3 min 2 s for 1896886 feature vectors) and evaluation is done in 23 s for 2312432 feature vectors.” Further Sarwar, page 3, col 1, section 4.1, ¶1 “As a result, the class imbalance problem is less stressed and the classifier has to be trained with only 215053 data in total, which results in short training time.” Here, the single training iteration can be considered a defined number of training iterations), and
(iii) a training iteration of the defined number of training iterations comprises:
(a) generating the first training entry batch, wherein a training entry of the first training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the primary event (Sarwar, page 2, col 2, section 3.4, ¶2 “After splitting, we randomly take half of the sessions in the clickBuy file to create our local test set and solution file;” here, the randomly selected half of the sessions can be considered the randomly generated set, and the solution file can be considered the ground truth label, with the buying being the dependent event),
(b) updating parameters of the identifier machine learning node using the first training entry batch (Sarwar, page 4, col 1, ¶2 “In order to address the problems we performed resampling of data and classified sessions using AdaBoost.M1 algorithm… For Adaboost.M1 the model is built in a smaller period of time (3 min 2 s for 1896886 feature vectors) and evaluation is done in 23 s for 2312432 feature vectors.”),
(c) generating a second training entry batch, wherein a training entry of the second training entry batch (1) is randomly generated from the set of filtered training entries and (2) comprises a ground-truth label corresponding to an occurrence of the dependent event (Sarwar, page 3, col 1, section 4.1, ¶1 “An important observation about item classifier is that we have to train the classifier only with the click data of the buy sessions in training data. The click data of a buy session contain a set of items that was bought (Bs) and a set of items that was not bought (As). Now for each item i ∈ Bs we extract both session-based and item-based features. After that we label that item as buy and each item i ∈ As as non buy. By considering only the buy sessions we get 1049817 bought items and 1264870 non-bought ones.”), and
(d) updating parameters of the differentiator machine learning node using the second training entry batch (Sarwar, page 3, col 1, section 4.1, ¶1 “As a result, the class imbalance problem is less stressed and the classifier has to be trained with only 215053 data in total, which results in short training time.”).
Kumar in view of Dong in view of Zhao in further view of Colucci and Sarwar are analogous art because both references concern methods for predicting outcomes and issues with dependent and independent events. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kumar/Dong/Zhao/Colucci’s method to incorporate two incorporated classifiers as taught by Sarwar. The motivation for doing so would have been to take advantage of Sarwar’s two classifiers to battle class imbalance as stated in Sarwar, page “The usage of a cascade of two different classifiers is beneficial when one experiences severe imbalance between buy and non-buy sessions and multiplicity of good feature subsets.”
Response to Arguments
Applicant's arguments filed February 23rd, 2026 (hereinafter “Remarks”) have been fully considered but they are not persuasive.
Regarding the objections to the Specification, Applicant’s amended Specification has overcome the objections, which are withdrawn.
Applicant’s arguments with respect to the 35 U.S.C. § 103 rejections have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Rejections under 35 U.S.C. § 101:
Argument 1:
“See Desjardins Memorandum, page 4. For at least this reason, the claims show an improvement in computer functionality that integrates any abstract idea into a practical application.” (Remarks, page 3).
Examiners Response:
Examiner respectfully disagrees, the MPEP states “Conversely, if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology.” See MPEP § 2106.04(d)(1). Further, the applicant merely uses a computer to perform processes which can be performed by a mental process. An improvement to balancing an imbalanced training data set may be an improvement in an abstract idea, but not an improvement in the functioning of a computer, as a computer. Unlike Desjardins, the additional elements are recited at a high level of generality, and even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application
Argument 2:
“Like the claims of Hannun, claim 1 recites a specific implementation of machine learning model calls that implement a particular machine learning pipeline which obtains a predictive output from a machine learning model, identifies candidate imbalance adjustment conditions, selects optimal imbalance adjustment conditions, determines a set of filtered training entries based on those conditions, and trains the machine learning model using the filtered training entries to improve the model's performance… Each of the above steps leverage a specific machine learning framework, which is not practically performed within the human mind.” (Remarks, pages 4-5).
Examiners Response:
Examiner respectfully disagrees, the MPEP states “Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind." Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015). See also Intellectual Ventures I LLC v. Symantec Corp., 838 F.3d 1307, 1318, 120 USPQ2d 1353, 1360 (Fed. Cir. 2016) (‘‘[W]ith the exception of generic computer-implemented steps, there is nothing in the claims themselves that foreclose them from being performed by a human, mentally or with pen and paper.’’); Mortgage Grader, Inc. v. First Choice Loan Servs. Inc., 811 F.3d 1314, 1324, 117 USPQ2d 1693, 1699 (Fed. Cir. 2016) (holding that computer-implemented method for "anonymous loan shopping" was an abstract idea because it could be "performed by humans without a computer"). Mental processes recited in claims that require computers are explained further below with respect to point C.” See MPEP § 2106.04(a)(2)(III). The identification, selection and determining steps all fall within the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Further elements such as the use of a machine learning model link the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
Argument 3:
“Claim 1 recites an improvement to machine learning training, which has been affirmatively recognized as an improvement sufficient to integrate a juridical limitation into a practical application. See Ex Parte Desjardines. Thus, even if claim 1 were directed to an abstract idea-which, Applicant submits, it is not the claim recites a combination of additional elements that improves a technical field such that the claim as a whole integrates any alleged abstract idea into a practical application” (Remarks, page 5).
Examiners Response:
Examiner respectfully disagrees, Ex Parte Desjardins did not recognize that all training limitations are an improvement to how the machine learning model itself operates, but instead the particular claims within Ex Parte Desjardins were found to be allowable. The MPEP states “To show that the involvement of a computer assists in improving the technology, the claims must recite the details regarding how a computer aids the method, the extent to which the computer aids the method, or the significance of a computer to the performance of the method. Merely adding generic computer components to perform the method is not sufficient. Thus, the claim must include more than mere instructions to perform the method on a generic component or machinery to qualify as an improvement to an existing technology.” See MPEP § 2106.05(a)(II). Here, the use of a computer does not integrate the abstract idea into a practical application, but instead can be considered an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Argument 4:
“By utilizing the techniques described in relation to various embodiments of the present invention, an imbalanced training set can nevertheless be used to train an effective and reliable machine learning model, a capability that avoids need to perform a large number of training iterations with large training sets to train suitable machine learning models, thus improving resource usage efficiency and operational throughput of computing systems that use classification machine learning models that are trained with imbalanced training sets.” (Remarks, page 7).
Examiners Response:
Examiner respectfully disagrees, the MPEP states “…"claiming the improved speed or efficiency inherent with applying the abstract idea on a computer" does not integrate a judicial exception into a practical application or provide an inventive concept. Intellectual Ventures I LLC v. Capital One Bank (USA), 792 F.3d 1363, 1367, 115 USPQ2d 1636, 1639 (Fed. Cir. 2015).” See MPEP § 2106.5(f). The use of generic computing components to perform the abstract idea does not integrate a judicial exception into a practical application.
Argument 5:
“In Ex Parte Desjardines, the Panel expressly recognized machine learning training as constituting "an improvement to how the machine learning model itself operations, and not, for example, the identified mathematical calculation." Ex Parte Desjardines, page 9. Therefore, Applicant respectfully submits that independent claim 1 recites patent eligible subject matter under 35 U.S.C. § 101 and requests withdrawal of the rejection to claim 1 (and the claims that depend therefrom).” (Remarks, page 7).
Examiners Response:
Examiner respectfully disagrees, Ex Parte Desjardins did not recognize that all training limitations are an improvement to how the machine learning model itself operates, but instead the particular claims within Ex Parte Desjardins were found to be allowable. Like in claim 2 of example 47 of the Al-related SME examples 47-49 issued in 2024, training limitations recited at a high level of generality, and using generic computing components, can be considered to fall under both an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). As well as generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h). Further, the additional elements are recited at a high level of generality, and even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application. Therefore, the claims are rejected under 35 U.S.C. § 101.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Al-Kofahi et al. (US 7,062,498 B2) discloses systems, methods, and software to aid classification of text, such as headnotes and other documents, to target classes in a target classification system. For example, one system computes composite scores based on: similarity of input text to text assigned to each of the target classes; similarity of non-target classes assigned to the input text and target classes; probability of a target class given a set of one or more non-target classes assigned to the input text; and/or probability of the input text given text assigned to the target classes. The exemplary system then evaluates the composite scores using class-specific decision criteria, such as thresholds, ultimately assigning or recommending assignment of the input text to one or more of the target classes.
Dong et al. (US 6,507,843 B1) discloses each EP can sharply differentiate the class membership of a (possibly small) fraction of instances containing the EP, due to the big difference between the EP's supports in the opposing classes; the differentiating power of the EP is defined in terms of the EP's supports and ratio, on instances containing the EP.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB Z SUSSMAN MOSS whose telephone number is (571) 272-1579. The examiner can normally be reached Monday - Friday, 9 a.m. - 5 p.m. ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.S.M./Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122