DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because FIG. 1G includes the reference character R not mentioned in the description. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-7 are directed to a process. Claims 8-20 are directed to a machine or an article of manufacture.
With respect to claim(s) 1:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
removing/remove […] protected dimensions from the training data to generate modified training data; (Mental process – A person can remove protected dimensions from training data to generate modified training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
[…] generate first predictions based on test data derived from the modified training data; (Mental process – A person can mentally generate predictions based on data – see MPEP § 2106.04(a)(2)(III))
[…] generate second predictions based on the test data; (Mental process – A person can mentally generate predictions based on data – see MPEP § 2106.04(a)(2)(III))
determining […] whether correlations between the first predictions and the second predictions are less than a threshold; (Mental process – A person can mentally determine correlations between predictions are less than a threshold – see MPEP § 2106.04(a)(2)(III))
selectively: determining […] that the first trained model is bias-reduced based on the correlations being less than the threshold; or determining […] that the first trained model is biased based on the correlations not being less than the threshold. (Mental process – A person can mentally determine whether correlations are less than or not less than a threshold – see MPEP § 2106.04(a)(2)(III))
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then the claim limitations fall within the mathematical or mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
A method, comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receiving, by a device, training data, a first model, and a second model; (Mere data gathering – Adding insignificant extra-solution activity of mere data gathering to the judicial exception – see § MPEP2106.05(g).)
training, by the device, the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
training, by the device, the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilizing, by the device, the first trained model to […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilizing, by the device, the second trained model to […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
[…] by the device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP § 2106.05(f).)
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A method, comprising: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receiving, by a device, training data, a first model, and a second model; (Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception (WURC)- see MPEP § 2106.05(d)(ll)(i) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information).)
training, by the device, the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
training, by the device, the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilizing, by the device, the first trained model to […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilizing, by the device, the second trained model to […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
[…] by the device […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP § 2106.05(f).)
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
With respect to claim(s) 2 and 16:
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 16) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
implementing/implement the first trained model based on determining that the first trained model is bias-reduced. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 16) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
implementing/implement the first trained model based on determining that the first trained model is bias-reduced. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 3 and 17:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
(Claim 3) determining that the first trained model is generating biased predictions, (Mental process – A person can mentally determine that a model is generating biased predictions – see MPEP § 2106.04(a)(2)(III))
(Claim 17) […] determining that the first trained model is biased. (Mental process – A person can mentally determine that a model is biased – see MPEP § 2106.04(a)(2)(III))
(Claim 3) wherein the retraining includes removing dimensions that are correlated with the protected dimensions. (Mental process – A person can remove dimensions that are correlated with protected dimensions via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 17) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
retraining/retrain the first trained model based on […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 17) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
retraining/retrain the first trained model based on […] (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 4 and 18:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
identifying/identify one or more additional protected dimensions that cause the correlations to not be less than the threshold; and removing/remove the one or more additional protected dimensions from the training data. (Mental process – A person can identify and remove protected dimensions that cause correlations to not be less than a threshold via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 18) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 18) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 5:
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
wherein the training data includes a plurality of dimensions and the protected dimensions include one or more dimensions associated with historical bias. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the training data includes a plurality of dimensions and the protected dimensions include one or more dimensions associated with historical bias. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 6 and 14:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
wherein the test data is derived from the modified training data by excluding the protected dimensions from the modified training data. (Mental process – A person can mentally exclude protected dimensions from the modified training data to derive test data – see MPEP § 2106.04(a)(2)(III))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 7 and 19:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
removing/remove secondary correlated protected dimensions from the modified training data prior to training the first model with the modified training data. (Mental process – A person can remove secondary correlated protected dimensions from the modified training data mentally or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
(Claim 19) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
(Claim 19) wherein the one or more instructions further cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 8:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
remove protected dimensions from the training data to generate modified training data; (Mental process – A person can remove protected dimensions from training data to generate modified training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
[…] generate first predictions based on test data derived from the modified training data; (Mental process – A person can mentally generate predictions based on data – see MPEP § 2106.04(a)(2)(III))
[…] generate second predictions based on the test data; (Mental process – A person can mentally generate predictions based on data – see MPEP § 2106.04(a)(2)(III))
determine whether correlations between the first predictions and the second predictions are less than a threshold; (Mental process – A person can mentally determine correlations between predictions are less than a threshold – see MPEP § 2106.04(a)(2)(III))
determine that the first trained model is bias-reduced based on the correlations being less than the threshold; (Mental process – A person can mentally determine whether correlations are less than a threshold – see MPEP § 2106.04(a)(2)(III))
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then the claim limitations fall within the mathematical or mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
A device, comprising: one or more processors configured to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receive training data, a first model, and a second model; (Mere data gathering – Adding insignificant extra-solution activity of mere data gathering to the judicial exception – see § MPEP2106.05(g).)
train the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
train the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the first trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the second trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
implement the first trained model based on determining that the first trained model is bias-reduced. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A device, comprising: one or more processors configured to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receive training data, a first model, and a second model; (Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception (WURC)- see MPEP § 2106.05(d)(ll)(i) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information).)
train the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
train the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the first trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the second trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
implement the first trained model based on determining that the first trained model is bias-reduced. (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
With respect to claim(s) 9 and 20:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
add cross-augmented dimensions to the modified training data prior to training the first model with the modified training data. (Mental process – A person can mentally add cross-augmented dimensions to modified training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 10:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
add cross-augmented data to replace the protected dimensions prior to training the first model with the modified training data, (Mental process – A person can mentally add cross-augmented data to modified training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
wherein the cross-augmented data includes values consistent with an original distribution of the protected dimensions. (Mathematical concepts – Cross-augmented data including values consistent with an original distribution recites a mathematical relationship (see paragraph [0035]) – see MPEP § 2106.04(a)(2)(I))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 11:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
remove trend-based information, related to the protected dimensions, from the modified training data prior to training the first model with the modified training data. (Mental process – A person can remove trend-based information from training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 12:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
adjust the threshold prior to determining whether the correlations between the first predictions and the second predictions are less than the threshold. (Mental process – A person can adjust a threshold via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 13:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
analyze an impact of removing the protected dimensions on an accuracy of the first model. (Mental process – A person can mentally analyze an impact of removing the protected dimensions on an accuracy of the model – see MPEP § 2106.04(a)(2)(III))
Additionally, the claim(s) do not recite any new additional elements that would amount to an integration of the abstract idea into a practical application (individually or in combination) or significantly more than the judicial exception.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Therefore, the claim is not patent eligible.
With respect to claim(s) 15:
2A Prong 1: The claim(s) recite(s) an abstract idea. Specifically:
remove protected dimensions from the training data to generate modified training data; (Mental process – A person can remove protected dimensions from training data to generate modified training data via mind or by using pen and paper – see MPEP § 2106.04(a)(2)(III))
[…] generate first predictions based on test data derived from the modified training data, wherein the test data is derived from the modified training data by excluding the protected dimensions from the modified training data; (Mental process – A person can mentally generate predictions based on data and can mentally generate test data by excluding protected dimensions from the data – see MPEP § 2106.04(a)(2)(III))
[…] generate second predictions based on the test data; (Mental process – A person can mentally generate predictions based on data – see MPEP § 2106.04(a)(2)(III))
determine whether correlations between the first predictions and the second predictions are less than a threshold; (Mental process – A person can mentally determine correlations between predictions are less than a threshold – see MPEP § 2106.04(a)(2)(III))
selectively: determine that the first trained model is bias-reduced based on the correlations being less than the threshold; or determine that the first trained model is biased based on the correlations not being less than the threshold. (Mental process – A person can mentally determine whether correlations are less than or not less than a threshold – see MPEP § 2106.04(a)(2)(III))
If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then the claim limitations fall within the mathematical or mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
2A Prong 2: The additional elements recited in the claim(s) do not integrate the abstract idea into a practical application, individually or in combination.
Additional elements:
A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device, cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receive training data, a first model, and a second model; (Mere data gathering – Adding insignificant extra-solution activity of mere data gathering to the judicial exception – see § MPEP2106.05(g).)
train the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
train the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the first trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the second trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea.
2B: The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device, cause the device to: (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
receive training data, a first model, and a second model; (Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception (WURC)- see MPEP § 2106.05(d)(ll)(i) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information).)
train the first model with the modified training data to generate a first trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
train the second model with the training data to generate a second trained model; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the first trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
utilize the second trained model to […]; (Mere instructions to apply an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-8 and 12-19 are rejected under 35 U.S.C. 103 as being unpatentable over VERMA ("Removing biased data to improve fairness and accuracy") in view of LAM (US 20230394366 A1) and WISNIEWSKI ("fairmodels: a Flexible Tool for Bias Detection, Visualization, and Mitigation in Binary Classification Models"), hereafter VERMA, LAM, and WISNIEWSKI, respectively.
Regarding Claim 1:
VERMA teaches:
A method, comprising: receiving […] training data, a first model, and a second model; (VERMA [page 4, Algorithm 2] teaches a Training dataset D as input (i.e., receiving […] training data). VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […]. For all experiments, the machine learning model architecture we used was a neural network with 2 hidden layers.” VERMA [page 8, section 4.4 Results for input space threshold
λ
=
0
.] teaches: “Figure 2 shows the remaining individual discrimination for all 240 models we trained for each experiment and each technique (i.e., a method) […].” Examiner’s note: VERMA trains 240 models for each experiment and each technique, each model being a neural network with 2 hidden layers. Therefore, under BRI, a first model can be interpreted as the model being trained on the dataset pre-processed via simple removal of the sensitive attribute (SR), and a second model can be interpreted as the model being trained on the full dataset. Further, the models must be retrieved, downloaded, initialized, instantiated, or similar prior to conducting any type of training. Thus, receiving […], a first model, and a second model can be interpreted as the step of acquiring the model prior to beginning training of the models.)
removing […] protected dimensions from the training data to generate modified training data; (VERMA [page 1, section 1 Introduction] teaches: “Bias in machine learning systems is undesirable because it can produce unfair decisions [23, 24] […]. Laws define sensitive attributes (e.g., race, sex, religion) that are illegal to use as the basis of any decision. […] We performed 8 experiments using 6 real-world datasets (some datasets contain more than one sensitive attribute).” VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […].” Examiner’s note: VERMA [page 1, section 1 Introduction] and [page 5, section 4 Evaluation] teaches pre-processing a dataset for training neural networks by removing sensitive attributes such as race, sex, religion. The protected dimensions can be interpreted as the sensitive attributes that are removed from the dataset (i.e., removing […] protected dimensions from the training data) when applying the simple removal processing technique to the dataset, thereby generating a dataset with the sensitive attributes removed (i.e., to generate modified training data).
training […] the first model with the modified training data to generate a first trained model; (VERMA, [page 2, section 2 Motivating Example] teaches: “Suppose that a bank wishes to automate the process of loan approval. The bank could train a machine learning model on historical loan decisions, then use the model to improve speed, reduce costs, and reduce subjectivity.” VERMA [page 3, Remove biased datapoints] teaches: “The bank should train its model on the fair decisions in the debiased dataset, rather than on the whole dataset. […] After identifying and removing the biased decisions, the bank can train their model by omitting the sensitive feature(s) (race in this case) to avoid disparate treatment [52]. Our approach empowers the bank to make less biased decisions in the future. (As with any automated decision-making process, the bank should include avenues for challenge and redress.) This avoids violating anti-discrimination laws and is the morally right thing to do.” VERMA [page 6, section 4.4 Results for input space threshold
λ
=
0
] teaches: “Answer to RQ2: SR also achieves 0% individual discrimination. This follows from our choice of input space similarity condition. When the sensitive feature is removed, the remaining features are the same for all pairs of similar individuals. And therefore, SR gives the same prediction for two individuals with the same features.” Examiner’s note: VERMA teaches training neural networks with two hidden layers using a dataset that has been pre-processed by removing or omitting the sensitive attributes by simple removal (SR) (i.e., training […] the first model with the modified data). The resulting neural network after training corresponds to a first trained model.)
training […] the second model with the training data to generate a second trained model; (VERMA [page 5, section 4 Evaluation] teaches: “a baseline model trained on the full training dataset (Full) […].”)
utilizing […] the first trained model to generate first predictions based on test data […] (VERMA [page 5, section 4.1 Experimental methodology] teaches: “Measure the model’s test accuracy on a debiased test set […].” VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Our evaluation methodology uses a debiased test set from which unfair points have been removed. The reason is that a user’s goal is not to obtain a model that performs well on the entire dataset that includes biased decisions, but a model that performs well on fair decisions. Our experimental results indicate that our debiasing. “VERMA [page 11, section 6 conclusion] teaches: “Compared to a baseline model that is trained on historical data without removing any datapoints, our technique improves test accuracy, individual discrimination, and statistical disparity. Ours is the only technique (out of eight tested) that improves all three measures, no matter which is chosen as the optimization goal.” Examiner’s note: VERMA trains and tests 240 models with various hyperparameter configurations, including a baseline model and a model trained on a dataset pre-processed to remove sensitive attributes, as disclosed in VERMA [page 5, section 4, Evaluation]. Therefore, under BRI, utilizing […] the first trained model to generate first predictions based on test data can be interpreted as testing the model trained on the dataset with sensitive attributes removed, which is tested using a debiased test set.)
utilizing […] the second trained model to generate second predictions based on the test data; (VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Another advantage is that using the same test set for a particular set of hyperparameters provides an apples-to-apples comparison of our technique with all the seven baselines. For the Full baseline (i.e., utilizing […] the second trained model), the full training set is used, while the test set (i.e., to generate second predictions based on the test data) is still debiased.”)
VERMA is not relied upon for teaching:
[…] by the device […]
[…] test data derived from the modified training data;
determining, by the device, whether correlations between the first predictions and the second predictions are less than a threshold; and
selectively: determining, by the device, that the first trained model is bias-reduced based on the correlations being less than the threshold; or determining, by the device, that the first trained model is biased based on the correlations not being less than the threshold.
However, LAM teaches: […] by the device […] (LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
[…] test data derived from the modified training data; (LAM [0027] teaches: “According to various embodiments, a protected attribute may be any feature for which bias is to be removed.” LAM [0056] teaches: “According to various embodiments, the default protected attribute values may be used during the test and inference phases to replace actual protected attribute values. Various approaches may be used to determine default protected attribute values. For example, protected attribute values may be dropped completely and treated as missing. As another example, protected attribute values may be replaced with a single value for all observations. For instance, in a data set in which each observation corresponds to a person, the race of each individual may be set to a default value (e.g., Black, White, etc.), while the gender of each individual may be set to a default value (e.g., female, male, etc.). In this way, the actual race and gender of an individual may be masked during the test and inference phases so that it may not generate disparate treatment bias.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610.” LAM [0085] teaches: “Test data for analysis is determined at 704. According to various embodiments, a training data set may be divided into data used to actively train the model and data used to test the performance of the training.” LAM [0086] teaches: “In some implementations, a test data set may be preprocessed using some or all of the techniques discussed with respect to operation 406 and the method 800 shown in FIG. 8. That is, the same rules used to determine default data values for feature values exhibiting insufficient overlap or positivity violations may be applied to the test data set.” Examiner’s note: LAM [0071] and [0085-0086] teach applying rules for removing features from the training data, and that these same rules may be applied for removing features from the test data set. Applying the same rules, such as removing race or gender from the training dataset, would result in removing the same protected attributes in the test data after dividing the initial training dataset. Therefore, under BRI, test data derived from the modified training data can be interpreted as a test dataset that is the result of applying the same protected attribute removal rules as its corresponding training dataset.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA and LAM before them, to include LAM’s computing system and rules for removing protected attribute from test data as in training data in VERMA’s bias removal method. One would have been motivated to make such a combination in order to apply same rules for removing protected attributes from test data similarly as in the training data so that the model may not generate disparate treatment bias (LAM [0056]).
VERMA in view of LAM is not relied upon for teaching, but WISNIEWSKI teaches: determining […] whether correlations between the first predictions and the second predictions are less than a threshold; (WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule. For example, for the metric Statistical Parity (STP), a model would be
ε
-non-discriminatory for privileged subgroup
a
if it satisfies
∀
b
∈
A
{
a
}
ε
<
S
T
P
r
a
t
i
o
=
S
T
P
b
S
T
P
a
<
1
ε
1
WISNIEWSKI [page 10, Visualizing bias] teaches: “So for example when we would like to know the parity loss of Statistical Parity between unprivileged (b) and privileged (a) subgroups we mean value like this:
S
T
P
p
a
r
i
t
y
_
l
o
s
s
=
ln
S
T
P
b
S
T
P
a
2
"
WISNIEWSKI [page 10, Table 2] teaches that
S
T
P
is defined as the positive rate in the form of
T
P
+
F
P
T
P
+
F
P
+
T
N
+
F
N
.” Examiner’s note: WISNIEWSKI teaches computing the positive rate as
S
T
P
a
for the privileged subgroup, and
S
T
P
b
for the unprivileged subgroup, using TP, FP, TN, and FN, which represent the model predictions. Under BRI, determining […] whether correlations between the first predictions and the second predictions are less than a threshold can be interpreted as computing
S
T
P
p
a
r
i
t
y
_
l
o
s
s
and determining whether it is below the acceptable and common epsilon choice of 0.8.)
selectively: determining […] that the first trained model is bias-reduced based on the correlations being less than the threshold; or determining, by the device, that the first trained model is biased based on the correlations not being less than the threshold. (WISNIEWSKI [page 2, Related Work] teaches: “The package fairmodels not only allows for that comparison between models and multiple exposed groups of people, but it gives direct feedback if the model is fair or not […].” WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, and WISNIEWSKI before them, to include WISNIEWSKI’s STP computation for determining the acceptable amount of bias in models used in VERMA and LAM’s bias removal method. One would have been motivated to make such a combination in order to implement a metric that measures the ratio between privileged and unprivileged subgroup to ensure an acceptable amount of bias of a model (WISNIEWSKI [page 4, Acceptable amount of bias]).
Regarding Claim 2:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. LAM further teaches:
implementing the first trained model based on determining that the first trained model is bias-reduced. (LAM [0017] teaches: “One or more performance metrics may be determined and used to evaluate the trained model and improve the training process. Then, to predict one or more unobserved outcome values, the trained prediction model may be applied to inference data.” LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0060] teaches: “The supervised machine learning model is stored on a storage device at 416. In some implementations, storing the supervised machine learning model may involve storing one or more weights or values suitable for use in applying the supervised machine learning model to novel data.” LAM [0078] teaches: “At 614, a determination is made as to whether the overlap values exceed a designated threshold. According to various embodiments, the designated threshold may be determined so as to avoid or prevent positivity violations, and may depend on the goals or context associated with the prediction model. For example, higher threshold levels may improve model prediction at the expense of increasing potential bias, while lower threshold levels may reduce potential bias but also model predictive power.” LAM [0079] teaches: “If one or more of the overlap values fail to exceed a designated threshold, then at 616 the feature values having insufficient overlap are replaced with default feature values. According to various embodiments, various approaches may be used to determine default feature values. For example, feature values with insufficient overlap may be dropped completely and treated as missing.” Examiner’s note: LAM discloses machine learning model training and evaluation methods for reducing bias prior to storing the model weights for inference (i.e., implementing the first trained model). Further, LAM [0060] teaches that the weights of the supervised machine learning model, which was bias reduced according to LAM [0024], are suitable for use with novel data. The suitability is determined based on the designated threshold disclosed in LAM [0078] and [0079]. Therefore, the model is applied to inference data based on determining that a bias-reduction threshold is satisfied (i.e., based on determining that the first trained model is bias-reduced).)
Regarding Claim 3:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. VERMA further teaches:
retraining the first trained model based on determining that the first trained model is generating biased predictions, (VERMA [page 4, section 3 Algorithm] teaches: “Given a trained model and a set of datapoints along with their predictions from the model, RankByInfluence ranks the training data of the model from most influential to least influential training datapoint responsible for those predictions. If the most influential datapoints are removed from the training data, and the model is retrained with the same model architecture, the probability of change in the prediction of the discriminatory pairs is highest. […] It then iteratively removes a chunk of the most biased datapoints from the sorted original training dataset (line 6), retrains the same model architecture on the remaining datapoints, and estimates the individual discrimination of the retrained model (line 7).”)
LAM further teaches: wherein the retraining includes removing dimensions that are correlated with the protected dimensions. (LAM [0039] teaches: “For instance, features that have low predictive power but that are highly correlated with values corresponding with the protected attribute A 302 may be automatically removed.”)
Regarding Claim 4:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. WISNIEWSKI further teaches:
identifying one or more additional protected dimensions that cause the correlations to not be less than the threshold; (WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule. For example, for the metric Statistical Parity (STP), a model would be
ε
-non-discriminatory for privileged subgroup
a
if it satisfies
∀
b
∈
A
{
a
}
ε
<
S
T
P
r
a
t
i
o
=
S
T
P
b
S
T
P
a
<
1
ε
1
WISNIEWSKI [page 2, Introduction] teaches: “Later in this paper, we present tools to identify differences between groups defined by some protected attribute but note that this does not automatically mean that there is discrimination.” WISNIEWSKI [page 3, Fairness metrics] teaches: “Machine learning models just like human-based decisions can be biased against people with certain sensitive attributes which are also called protected groups. They consist of subgroups - people who share the same sensitive attribute, like gender, race, or some other feature.” WISNIEWSKI [page 6, Figure 2] teaches: “The dots (unprivileged subgroups - female) and vertical lines (privileged subgroup - male) of the metric scores plot represent raw metric scores of subgroups.” WISNIEWSKI [page 10, Visualizing bias] teaches: “So for example when we would like to know the parity loss of Statistical Parity between unprivileged (b) and privileged (a) subgroups we mean value like this:
S
T
P
p
a
r
i
t
y
_
l
o
s
s
=
ln
S
T
P
b
S
T
P
a
2
"
WISNIEWSKI [page 10, Table 2] teaches that
S
T
P
is defined as the positive rate in the form of
T
P
+
F
P
T
P
+
F
P
+
T
N
+
F
N
.” Examiner’s note: WISNIEWSKI teaches groups of protected attributes such as gender, race, or some other feature, and subgroups for each protected attribute like male and female for gender, for example. Each
S
T
P
r
a
t
i
o
is calculated for subgroups of the same protected attribute group. Under BRI, identifying one or more additional protected dimensions that cause the correlations to not be less than the threshold can be interpreted as when a computed
S
T
P
r
a
t
i
o
is not within the interval
ε
,
1
ε
.)
LAM further teaches: removing the one or more additional protected dimensions from the training data. (LAM [0096] teaches: “In some embodiments, the operation 808 may involve entirely removing data values corresponding with the protected attribute. For instance, a sex or race parameter in a supervised machine learning model may be dropped entirely.”)
Regarding Claim 5:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. VERMA further teaches:
wherein the training data includes a plurality of dimensions and the protected dimensions include one or more dimensions associated with historical bias. (VERMA [page 1, section 1 Introduction] teaches: “For example, loan applications and job applications from minority communities have been more frequently denied (i.e., associated with historical bias) [42, 103, 115]. Training on historical data would perpetuate that injustice. Selection bias occurs when selecting samples from a demographic group inadvertently introduces undesired correlations between the features (i.e., a plurality of dimensions) pertaining to that demographic group and the training labels [12, 113, 120], e.g. in the selected subsample for a group, most of their loan requests were denied.” VERMA [pages 5-6, section 4.2 Experiments with input space threshold
λ
=
0
] teaches: “For example, the Salary dataset has features sex, rank, age, degree, and experience, out of which age and experience are numerical while the others are categorical features.” VERMA [page 11, section 6 Conclusion] teaches: “Compared to a baseline model that is trained on historical data (i.e., the training data) without removing any datapoints (i.e., a plurality of dimensions and the protected dimensions include one or more dimensions associated with historical bias), our technique improves test accuracy, individual discrimination, and statistical disparity.”
Regarding Claim 6:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. LAM further teaches:
wherein the test data is derived from the modified training data by excluding the protected dimensions from the modified training data. (LAM [0027] teaches: “According to various embodiments, a protected attribute may be any feature for which bias is to be removed.” LAM [0056] teaches: “According to various embodiments, the default protected attribute values may be used during the test and inference phases to replace actual protected attribute values. Various approaches may be used to determine default protected attribute values. For example, protected attribute values may be dropped completely and treated as missing. As another example, protected attribute values may be replaced with a single value for all observations. For instance, in a data set in which each observation corresponds to a person, the race of each individual may be set to a default value (e.g., Black, White, etc.), while the gender of each individual may be set to a default value (e.g., female, male, etc.). In this way, the actual race and gender of an individual may be masked during the test and inference phases so that it may not generate disparate treatment bias.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610.” LAM [0085] teaches: “Test data for analysis is determined at 704. According to various embodiments, a training data set may be divided into data used to actively train the model and data used to test the performance of the training.” LAM [0086] teaches: “In some implementations, a test data set may be preprocessed using some or all of the techniques discussed with respect to operation 406 and the method 800 shown in FIG. 8. That is, the same rules used to determine default data values for feature values exhibiting insufficient overlap or positivity violations may be applied to the test data set.” Examiner’s note: LAM [0071] and [0085-0086] teach applying rules for removing features from the training data, and that these same rules may be applied for removing features from the test data set. Applying the same rules, such as removing race or gender from the training dataset, would result in removing the same protected attributes in the test data after dividing the initial training dataset. Therefore, under BRI, test data is derived from the modified training data can be interpreted as a test dataset that is the result of applying the same protected attribute removal rules (i.e., by excluding the protected dimensions from the modified training data) as its corresponding training dataset.)
Regarding Claim 7:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 1 as outlined above. LAM further teaches:
removing secondary correlated protected dimensions from the modified training data prior to training the first model with the modified training data. (LAM [0039] teaches: “For instance, features that have low predictive power but that are highly correlated with values corresponding with the protected attribute A 302 (i.e., secondary correlated protected dimensions) may be automatically removed.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610 (i.e., removing from the modified training data).” LAM [0039] teaches: “In FIG. 3B, protected attribute pure proxies A′ 312 purely proxies for the protected attribute A 302 and are removed from the set of features used in both training (i.e., prior to training the first model with the modified training data) and inference.”)
Regarding Claim 8:
VERMA teaches:
receive training data, a first model, and a second model; (VERMA [page 4, Algorithm 2] teaches a Training dataset D as input (i.e., receive training data). VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […]. For all experiments, the machine learning model architecture we used was a neural network with 2 hidden layers.” VERMA [page 8, section 4.4 Results for input space threshold
λ
=
0
.] teaches: “Figure 2 shows the remaining individual discrimination for all 240 models we trained for each experiment and each technique […].” Examiner’s note: VERMA trains 240 models for each experiment and each technique, each model being a neural network with 2 hidden layers. Therefore, under BRI, a first model can be interpreted as the model being trained on the dataset pre-processed via simple removal of the sensitive attribute (SR), and a second model can be interpreted as the model being trained on the full dataset. Further, the models must be retrieved, downloaded, initialized, instantiated, or similar prior to conducting any type of training. Thus, receiving […], a first model, and a second model can be interpreted as the step of acquiring the models prior to beginning training of the models.)
remove protected dimensions from the training data to generate modified training data; (VERMA [page 1, section 1 Introduction] teaches: “Bias in machine learning systems is undesirable because it can produce unfair decisions [23, 24] […]. Laws define sensitive attributes (e.g., race, sex, religion) that are illegal to use as the basis of any decision. […] We performed 8 experiments using 6 real-world datasets (some datasets contain more than one sensitive attribute).” VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […].” Examiner’s note: VERMA [page 1, section 1 Introduction] and [page 5, section 4 Evaluation] teaches pre-processing a dataset for training neural networks by removing sensitive attributes such as race, sex, religion. The protected dimensions can be interpreted as the sensitive attributes that are removed from the dataset (i.e., remove protected dimensions from the training data) when applying the simple removal processing technique to the dataset, thereby generating a dataset with the sensitive attributes removed (i.e., to generate modified training data).
train the first model with the modified training data to generate a first trained model; (VERMA, [page 2, section 2 Motivating Example] teaches: “Suppose that a bank wishes to automate the process of loan approval. The bank could train a machine learning model on historical loan decisions, then use the model to improve speed, reduce costs, and reduce subjectivity.” VERMA [page 3, Remove biased datapoints] teaches: “The bank should train its model on the fair decisions in the debiased dataset, rather than on the whole dataset. […] After identifying and removing the biased decisions, the bank can train their model by omitting the sensitive feature(s) (race in this case) to avoid disparate treatment [52]. Our approach empowers the bank to make less biased decisions in the future. (As with any automated decision-making process, the bank should include avenues for challenge and redress.) This avoids violating anti-discrimination laws and is the morally right thing to do.” VERMA [page 6, section 4.4 Results for input space threshold
λ
=
0
] teaches: “Answer to RQ2: SR also achieves 0% individual discrimination. This follows from our choice of input space similarity condition. When the sensitive feature is removed, the remaining features are the same for all pairs of similar individuals. And therefore, SR gives the same prediction for two individuals with the same features.” Examiner’s note: VERMA teaches training neural networks with two hidden layers using a dataset that has been pre-processed by removing or omitting the sensitive attributes by simple removal (SR) (i.e., train the first model with the modified data). The resulting neural network after training corresponds to a first trained model.)
train the second model with the training data to generate a second trained model; (VERMA [page 5, section 4 Evaluation] teaches: “a baseline model trained on the full training dataset (Full) […].”)
utilize the first trained model to generate first predictions based on test data […]; (VERMA [page 5, section 4.1 Experimental methodology] teaches: “Measure the model’s test accuracy on a debiased test set […].” VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Our evaluation methodology uses a debiased test set from which unfair points have been removed. The reason is that a user’s goal is not to obtain a model that performs well on the entire dataset that includes biased decisions, but a model that performs well on fair decisions. Our experimental results indicate that our debiasing.” VERMA [page 11, section 6 conclusion] teaches: “Compared to a baseline model that is trained on historical data without removing any datapoints, our technique improves test accuracy, individual discrimination, and statistical disparity. Ours is the only technique (out of eight tested) that improves all three measures, no matter which is chosen as the optimization goal.” Examiner’s note: VERMA trains and tests 240 models with various hyperparameter configurations, including a baseline model and a model trained on a dataset pre-processed to remove sensitive attributes, as disclosed in VERMA [page 5, section 4, Evaluation]. Therefore, under BRI, utilize the first trained model to generate first predictions based on test data can be interpreted as testing the model trained on the dataset with sensitive attributes removed, which is tested using a debiased test set.)
utilize the second trained model to generate second predictions based on the test data; (VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Another advantage is that using the same test set for a particular set of hyperparameters provides an apples-to-apples comparison of our technique with all the seven baselines. For the Full baseline (i.e., utilize the second trained model), the full training set is used, while the test set (i.e., to generate second predictions based on the test data) is still debiased.”)
VERMA is not relied upon for teaching:
A device, comprising: one or more processors configured to:
[…] test data derived from the modified training data;
determine whether correlations between the first predictions and the second predictions are less than a threshold;
determine that the first trained model is bias-reduced based on the correlations being less than the threshold;
implement the first trained model based on determining that the first trained model is bias-reduced.
However, LAM teaches: A device, comprising: one or more processors configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
[…] test data derived from the modified training data; (LAM [0027] teaches: “According to various embodiments, a protected attribute may be any feature for which bias is to be removed.” LAM [0056] teaches: “According to various embodiments, the default protected attribute values may be used during the test and inference phases to replace actual protected attribute values. Various approaches may be used to determine default protected attribute values. For example, protected attribute values may be dropped completely and treated as missing. As another example, protected attribute values may be replaced with a single value for all observations. For instance, in a data set in which each observation corresponds to a person, the race of each individual may be set to a default value (e.g., Black, White, etc.), while the gender of each individual may be set to a default value (e.g., female, male, etc.). In this way, the actual race and gender of an individual may be masked during the test and inference phases so that it may not generate disparate treatment bias.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610.” LAM [0085] teaches: “Test data for analysis is determined at 704. According to various embodiments, a training data set may be divided into data used to actively train the model and data used to test the performance of the training.” LAM [0086] teaches: “In some implementations, a test data set may be preprocessed using some or all of the techniques discussed with respect to operation 406 and the method 800 shown in FIG. 8. That is, the same rules used to determine default data values for feature values exhibiting insufficient overlap or positivity violations may be applied to the test data set.” Examiner’s note: LAM [0071] and [0085-0086] teach applying rules for removing features from the training data, and that these same rules may be applied for removing features from the test data set. Applying the same rules, such as removing race or gender from the training dataset, would result in removing the same protected attributes in the test data after dividing the initial training dataset. Therefore, under BRI, test data derived from the modified training data can be interpreted as a test dataset that is the result of applying the same protected attribute removal rules as its corresponding training dataset.)
implement the first trained model based on determining that the first trained model is bias-reduced. (LAM [0017] teaches: “One or more performance metrics may be determined and used to evaluate the trained model and improve the training process. Then, to predict one or more unobserved outcome values, the trained prediction model may be applied to inference data.” LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0060] teaches: “The supervised machine learning model is stored on a storage device at 416. In some implementations, storing the supervised machine learning model may involve storing one or more weights or values suitable for use in applying the supervised machine learning model to novel data.” LAM [0078] teaches: “At 614, a determination is made as to whether the overlap values exceed a designated threshold. According to various embodiments, the designated threshold may be determined so as to avoid or prevent positivity violations, and may depend on the goals or context associated with the prediction model. For example, higher threshold levels may improve model prediction at the expense of increasing potential bias, while lower threshold levels may reduce potential bias but also model predictive power.” LAM [0079] teaches: “If one or more of the overlap values fail to exceed a designated threshold, then at 616 the feature values having insufficient overlap are replaced with default feature values. According to various embodiments, various approaches may be used to determine default feature values. For example, feature values with insufficient overlap may be dropped completely and treated as missing.” Examiner’s note: LAM discloses machine learning model training and evaluation methods for reducing bias prior to storing the model weights for inference (i.e., implement the first trained model). Further, LAM [0060] teaches that the weights of the supervised machine learning model, which was bias reduced according to LAM [0024], are suitable for use with novel data. The suitability is determined based on the designated threshold disclosed in LAM [0078] and [0079]. Therefore, the model is applied to inference data based on determining that a bias-reduction threshold is satisfied (i.e., based on determining that the first trained model is bias-reduced).)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA and LAM before them, to include LAM’s computing system, rules for removing protected attribute from test data as removed similarly from training data, and storing weights of the model for inference in VERMA’s bias removal method. One would have been motivated to make such a combination in order to apply same rules for removing protected attributes from test data similarly as in the training data so that the model may not generate disparate treatment bias (LAM [0056]).
VERMA in view of LAM is not relied upon for teaching, but WISNIEWSKI teaches: determine whether correlations between the first predictions and the second predictions are less than a threshold; (WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule. For example, for the metric Statistical Parity (STP), a model would be
ε
-non-discriminatory for privileged subgroup
a
if it satisfies
∀
b
∈
A
{
a
}
ε
<
S
T
P
r
a
t
i
o
=
S
T
P
b
S
T
P
a
<
1
ε
1
WISNIEWSKI [page 10, Visualizing bias] teaches: “So for example when we would like to know the parity loss of Statistical Parity between unprivileged (b) and privileged (a) subgroups we mean value like this:
S
T
P
p
a
r
i
t
y
_
l
o
s
s
=
ln
S
T
P
b
S
T
P
a
2
"
WISNIEWSKI [page 10, Table 2] teaches that
S
T
P
is defined as the positive rate in the form of
T
P
+
F
P
T
P
+
F
P
+
T
N
+
F
N
.” Examiner’s note: WISNIEWSKI teaches computing the positive rate as
S
T
P
a
for the privileged subgroup, and
S
T
P
b
for the unprivileged subgroup, using TP, FP, TN, and FN, which represent the model predictions. Under BRI, determining […] whether correlations between the first predictions and the second predictions are less than a threshold can be interpreted as computing
S
T
P
p
a
r
i
t
y
_
l
o
s
s
and determining whether it is below the acceptable and common epsilon choice of 0.8.)
determine that the first trained model is bias-reduced based on the correlations being less than the threshold; (WISNIEWSKI [page 2, Related Work] teaches: “The package fairmodels not only allows for that comparison between models and multiple exposed groups of people, but it gives direct feedback if the model is fair or not […].” WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, and WISNIEWSKI before them, to include WISNIEWSKI’s STP computation for determining the acceptable amount of bias in models used in VERMA and LAM’s bias removal method. One would have been motivated to make such a combination in order to implement a metric that measures the ratio between privileged and unprivileged subgroup to ensure an acceptable amount of bias of a model (WISNIEWSKI [page 4, Acceptable amount of bias]).
Regarding Claim 12:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. LAM further teaches:
wherein the one or more processors are further configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
WISNIEWSKI further teaches: adjust the threshold prior to determining whether the correlations between the first predictions and the second predictions are less than the threshold. (WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “In the implementation, this boundary is represented by
ε
and it is adjustable by the user but the default value will be 0.8. This rule is often used, but in each specific case, one should see if the fairness criteria should be set differently. Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule. For example, for the metric Statistical Parity (STP), a model would be
ε
-non-discriminatory for privileged subgroup
a
if it satisfies
∀
b
∈
A
{
a
}
ε
<
S
T
P
r
a
t
i
o
=
S
T
P
b
S
T
P
a
<
1
ε
1
WISNIEWSKI [page 10, Visualizing bias] teaches: “So for example when we would like to know the parity loss of Statistical Parity between unprivileged (b) and privileged (a) subgroups we mean value like this:
S
T
P
p
a
r
i
t
y
_
l
o
s
s
=
ln
S
T
P
b
S
T
P
a
2
"
Examiner’s note: WISNIEWSKI teaches computing the positive rate as
S
T
P
a
for the privileged subgroup, and
S
T
P
b
for the unprivileged subgroup, using TP, FP, TN, and FN, which represent the model predictions. Under BRI, determining […] whether correlations between the first predictions and the second predictions are less than a threshold can be interpreted as computing
S
T
P
p
a
r
i
t
y
_
l
o
s
s
and determining whether it is below the acceptable and common epsilon choice of 0.8.)
Regarding Claim 13:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. LAM further teaches:
wherein the one or more processors are further configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
VERMA further teaches: analyze an impact of removing the protected dimensions on an accuracy of the first model. (VERMA [page 6, section 4.3 Experiments with input space threshold
λ
>
0
] teaches: “Answer to RQ4: DIR, PS, and LFR degrade the accuracy; AD and MA affect it little; and SR improves it, though not as much as our technique does.” Examiner’s note: VERMA states that SR (simple removal of the sensitive attribute) improves the accuracy of the model.)
Regarding Claim 14:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale.
Regarding Claim 15:
VERMA teaches:
receive training data, a first model, and a second model; (VERMA [page 4, Algorithm 2] teaches a Training dataset D as input (i.e., receive training data). VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […]. For all experiments, the machine learning model architecture we used was a neural network with 2 hidden layers.” VERMA [page 8, section 4.4 Results for input space threshold
λ
=
0
.] teaches: “Figure 2 shows the remaining individual discrimination for all 240 models we trained for each experiment and each technique […].” Examiner’s note: VERMA trains 240 models for each experiment and each technique, each model being a neural network with 2 hidden layers. Therefore, under BRI, a first model can be interpreted as the model being trained on the dataset pre-processed via simple removal of the sensitive attribute (SR), and a second model can be interpreted as the model being trained on the full dataset. Further, the models must be retrieved, downloaded, initialized, instantiated, or similar prior to conducting any type of training. Thus, receiving […], a first model, and a second model can be interpreted as the step of acquiring the models prior to beginning training of the models.)
remove protected dimensions from the training data to generate modified training data; (VERMA [page 1, section 1 Introduction] teaches: “Bias in machine learning systems is undesirable because it can produce unfair decisions [23, 24] […]. Laws define sensitive attributes (e.g., race, sex, religion) that are illegal to use as the basis of any decision. […] We performed 8 experiments using 6 real-world datasets (some datasets contain more than one sensitive attribute).” VERMA [page 5, section 4 Evaluation] teaches: “We compared our pre-processing technique against 7 other techniques: a baseline model trained on the full training dataset (Full), five pre-processing techniques — simple removal of the sensitive attribute (SR), […].” Examiner’s note: VERMA [page 1, section 1 Introduction] and [page 5, section 4 Evaluation] teaches pre-processing a dataset for training neural networks by removing sensitive attributes such as race, sex, religion. The protected dimensions can be interpreted as the sensitive attributes that are removed from the dataset (i.e., remove protected dimensions from the training data) when applying the simple removal processing technique to the dataset, thereby generating a dataset with the sensitive attributes removed (i.e., to generate modified training data).
train the first model with the modified training data to generate a first trained model; (VERMA, [page 2, section 2 Motivating Example] teaches: “Suppose that a bank wishes to automate the process of loan approval. The bank could train a machine learning model on historical loan decisions, then use the model to improve speed, reduce costs, and reduce subjectivity.” VERMA [page 3, Remove biased datapoints] teaches: “The bank should train its model on the fair decisions in the debiased dataset, rather than on the whole dataset. […] After identifying and removing the biased decisions, the bank can train their model by omitting the sensitive feature(s) (race in this case) to avoid disparate treatment [52]. Our approach empowers the bank to make less biased decisions in the future. (As with any automated decision-making process, the bank should include avenues for challenge and redress.) This avoids violating anti-discrimination laws and is the morally right thing to do.” VERMA [page 6, section 4.4 Results for input space threshold
λ
=
0
] teaches: “Answer to RQ2: SR also achieves 0% individual discrimination. This follows from our choice of input space similarity condition. When the sensitive feature is removed, the remaining features are the same for all pairs of similar individuals. And therefore, SR gives the same prediction for two individuals with the same features.” Examiner’s note: VERMA teaches training neural networks with two hidden layers using a dataset that has been pre-processed by removing or omitting the sensitive attributes by simple removal (SR) (i.e., train the first model with the modified data). The resulting neural network after training corresponds to a first trained model.)
train the second model with the training data to generate a second trained model; (VERMA [page 5, section 4 Evaluation] teaches: “a baseline model trained on the full training dataset (Full) […].”)
utilize the first trained model to generate first predictions based on test data […]; (VERMA [page 5, section 4.1 Experimental methodology] teaches: “Measure the model’s test accuracy on a debiased test set […].” VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Our evaluation methodology uses a debiased test set from which unfair points have been removed. The reason is that a user’s goal is not to obtain a model that performs well on the entire dataset that includes biased decisions, but a model that performs well on fair decisions. Our experimental results indicate that our debiasing. “VERMA [page 11, section 6 conclusion] teaches: “Compared to a baseline model that is trained on historical data without removing any datapoints, our technique improves test accuracy, individual discrimination, and statistical disparity. Ours is the only technique (out of eight tested) that improves all three measures, no matter which is chosen as the optimization goal.” Examiner’s note: VERMA trains and tests 240 models with various hyperparameter configurations, including a baseline model and a model trained on a dataset pre-processed to remove sensitive attributes, as disclosed in VERMA [page 5, section 4, Evaluation]. Therefore, under BRI, utilize the first trained model to generate first predictions based on test data can be interpreted as testing the model trained on the dataset with sensitive attributes removed, which is tested using a debiased test set.)
utilize the second trained model to generate second predictions based on the test data; (VERMA [page 5, 4.1.1 Debiasing the test set.] teaches: “Another advantage is that using the same test set for a particular set of hyperparameters provides an apples-to-apples comparison of our technique with all the seven baselines. For the Full baseline (i.e., utilize the second trained model), the full training set is used, while the test set (i.e., to generate second predictions based on the test data) is still debiased.”)
VERMA is not relied upon for teaching:
A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:
one or more instructions that, when executed by one or more processors of a device, cause the device to:
[…] test data derived from the modified training data;
wherein the test data is derived from the modified training data by excluding the protected dimensions from the modified training data;
determine whether correlations between the first predictions and the second predictions are less than a threshold; and
selectively: determine that the first trained model is bias-reduced based on the correlations being less than the threshold; or determine that the first trained model is biased based on the correlations not being less than the threshold.
However, LAM teaches: A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device, cause the device to: (LAM [0109] teaches: “[0109] Any of the disclosed implementations may be embodied in various types of hardware, software, firmware, computer readable media, and combinations thereof. For example, some techniques disclosed herein may be implemented, at least in part, by computer-readable media that include program instructions, state information, etc., for configuring a computing system to perform various services and operations described herein. Examples of program instructions include both machine code, such as produced by a compiler, and higher-level code that may be executed via an interpreter. Instructions may be embodied in any suitable language such as, for example, Java, Python, C++, C, HTML, any other markup language, JavaScript, ActiveX, VBScript, or Perl. Examples of computer-readable media include, but are not limited to: magnetic media such as hard disks and magnetic tape; optical media such as flash memory, compact disk (CD) or digital versatile disk (DVD); magneto-optical media; and other hardware devices such as read-only memory (“ROM”) devices and random-access memory (“RAM”) devices. A computer-readable medium may be any combination of such storage devices.” LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
[…] test data derived from the modified training data; (LAM [0027] teaches: “According to various embodiments, a protected attribute may be any feature for which bias is to be removed.” LAM [0056] teaches: “According to various embodiments, the default protected attribute values may be used during the test and inference phases to replace actual protected attribute values. Various approaches may be used to determine default protected attribute values. For example, protected attribute values may be dropped completely and treated as missing. As another example, protected attribute values may be replaced with a single value for all observations. For instance, in a data set in which each observation corresponds to a person, the race of each individual may be set to a default value (e.g., Black, White, etc.), while the gender of each individual may be set to a default value (e.g., female, male, etc.). In this way, the actual race and gender of an individual may be masked during the test and inference phases so that it may not generate disparate treatment bias.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610.” LAM [0085] teaches: “Test data for analysis is determined at 704. According to various embodiments, a training data set may be divided into data used to actively train the model and data used to test the performance of the training.” LAM [0086] teaches: “In some implementations, a test data set may be preprocessed using some or all of the techniques discussed with respect to operation 406 and the method 800 shown in FIG. 8. That is, the same rules used to determine default data values for feature values exhibiting insufficient overlap or positivity violations may be applied to the test data set.” Examiner’s note: LAM [0071] and [0085-0086] teach applying rules for removing features from the training data, and that these same rules may be applied for removing features from the test data set. Applying the same rules, such as removing race or gender from the training dataset, would result in removing the same protected attributes in the test data after dividing the initial training dataset. Therefore, under BRI, test data derived from the modified training data can be interpreted as a test dataset that is the result of applying the same protected attribute removal rules as its corresponding training dataset.)
wherein the test data is derived from the modified training data by excluding the protected dimensions from the modified training data. (LAM [0027] teaches: “According to various embodiments, a protected attribute may be any feature for which bias is to be removed.” LAM [0056] teaches: “According to various embodiments, the default protected attribute values may be used during the test and inference phases to replace actual protected attribute values. Various approaches may be used to determine default protected attribute values. For example, protected attribute values may be dropped completely and treated as missing. As another example, protected attribute values may be replaced with a single value for all observations. For instance, in a data set in which each observation corresponds to a person, the race of each individual may be set to a default value (e.g., Black, White, etc.), while the gender of each individual may be set to a default value (e.g., female, male, etc.). In this way, the actual race and gender of an individual may be masked during the test and inference phases so that it may not generate disparate treatment bias.” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610.” LAM [0085] teaches: “Test data for analysis is determined at 704. According to various embodiments, a training data set may be divided into data used to actively train the model and data used to test the performance of the training.” LAM [0086] teaches: “In some implementations, a test data set may be preprocessed using some or all of the techniques discussed with respect to operation 406 and the method 800 shown in FIG. 8. That is, the same rules used to determine default data values for feature values exhibiting insufficient overlap or positivity violations may be applied to the test data set.” Examiner’s note: LAM [0071] and [0085-0086] teach applying rules for removing features from the training data, and that these same rules may be applied for removing features from the test data set. Applying the same rules, such as removing race or gender from the training dataset, would result in removing the same protected attributes in the test data after dividing the initial training dataset. Therefore, under BRI, test data is derived from the modified training data can be interpreted as a test dataset that is the result of applying the same protected attribute removal rules (i.e., by excluding the protected dimensions from the modified training data) as its corresponding training dataset.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA and LAM before them, to include LAM’s computing system and rules for removing protected attribute from test data as in training data in VERMA’s bias removal method. One would have been motivated to make such a combination in order to apply same rules for removing protected attributes from test data similarly as in the training data so that the model may not generate disparate treatment bias (LAM [0056]).
VERMA in view of LAM is not relied upon for teaching, but WISNIEWSKI teaches: determine whether correlations between the first predictions and the second predictions are less than a threshold; (WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule. For example, for the metric Statistical Parity (STP), a model would be
ε
-non-discriminatory for privileged subgroup
a
if it satisfies
∀
b
∈
A
{
a
}
ε
<
S
T
P
r
a
t
i
o
=
S
T
P
b
S
T
P
a
<
1
ε
1
WISNIEWSKI [page 10, Visualizing bias] teaches: “So for example when we would like to know the parity loss of Statistical Parity between unprivileged (b) and privileged (a) subgroups we mean value like this:
S
T
P
p
a
r
i
t
y
_
l
o
s
s
=
ln
S
T
P
b
S
T
P
a
2
"
WISNIEWSKI [page 10, Table 2] teaches that
S
T
P
is defined as the positive rate in the form of
T
P
+
F
P
T
P
+
F
P
+
T
N
+
F
N
.” Examiner’s note: WISNIEWSKI teaches computing the positive rate as
S
T
P
a
for the privileged subgroup, and
S
T
P
b
for the unprivileged subgroup, using TP, FP, TN, and FN, which represent the model predictions. Under BRI, determining […] whether correlations between the first predictions and the second predictions are less than a threshold can be interpreted as computing
S
T
P
p
a
r
i
t
y
_
l
o
s
s
and determining whether it is below the acceptable and common epsilon choice of 0.8.)
selectively: determine that the first trained model is bias-reduced based on the correlations being less than the threshold; or determine that the first trained model is biased based on the correlations not being less than the threshold. (WISNIEWSKI [page 2, Related Work] teaches: “The package fairmodels not only allows for that comparison between models and multiple exposed groups of people, but it gives direct feedback if the model is fair or not […].” WISNIEWSKI [page 4, Acceptable amount of bias] teaches: “Let
ε
>
0
(i.e., a threshold) be the acceptable amount of a bias. In this article we would say that the model is not discriminatory for a particular metric if the ratio between every unprivileged
b
,
c
, ... and privileged subgroup
a
is within
ε
,
1
ε
. The common choice for the epsilon is 0.8, which corresponds to the four-fifths rule.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, and WISNIEWSKI before them, to include WISNIEWSKI’s STP computation for determining the acceptable amount of bias in models used in VERMA and LAM’s bias removal method. One would have been motivated to make such a combination in order to implement a metric that measures the ratio between privileged and unprivileged subgroup to ensure an acceptable amount of bias of a model (WISNIEWSKI [page 4, Acceptable amount of bias]).
Regarding Claim 16:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 15 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Regarding Claim 17:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 15 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Regarding Claim 18:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 15 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Regarding Claim 19:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 15 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale.
Claims 9-10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over VERMA in view of LAM and WISNIEWSKI as applied respectively above to claims 8 and 15, and further in view of ROMANO ("Achieving Equalized Odds by Resampling Sensitive Attributes"), hereafter ROMANO.
Regarding Claim 9:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. LAM further teaches:
wherein the one or more processors are further configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
VERMA in view of LAM and WISNIEWSKI is not relied upon for teaching, but ROMANO teaches: add cross-augmented dimensions to the modified training data prior to training the first model with the modified training data. (ROMANO [page 1, section 1 Introduction] teaches: “Fairness constraints can often be articulated as conditional independence relations, and in this work we will focus on the equalized odds criterion [6], defined as
Y
^
A
|
Y
1
where the relationship above applies to test points; here,
Y
is the response variable,
A
is a sensitive attribute (e.g. gender),
X
is a vector of features that may also contain
A
, and
Y
^
=
f
^
X
is the prediction obtained with a fixed prediction rule
f
^
⋅
.” ROMANO [page 3, section 2.1 Regularization with fair dummies] teaches: “Our procedure starts by constructing a fair dummy sensitive attribute
A
~
i
for each training sample:
A
~
i
~
P
A
|
Y
A
i
|
Y
i
,
i
∈
I
t
r
a
i
n
,
where
P
A
|
Y
denotes the conditional distribution of
A
i
given
Y
i
. This sampling is straightforward; see (4) below. Importantly, we generate
A
~
i
without looking at
Y
^
i
so that we have the following property:
Y
^
i
A
~
i
|
Y
i
,
i
∈
I
t
r
a
i
n
,
2
Notice that the above is exactly the equalized odds relation in (1), with a crucial difference that the original sensitive attribute
A
i
is replaced by the artificial one
A
~
i
. […] we define
X
∈
R
I
t
r
a
i
n
×
p
,
A
~
∈
R
I
t
r
a
i
n
, and
Y
∈
R
I
t
r
a
i
n
, whose entries correspond to the features, sensitive attributes, fair dummies, and labels, respectively.” ROMANO [page 5, Algorithm 1] teaches step 2, which samples the fair dummies
A
~
i
~
P
A
|
Y
A
i
|
Y
i
and replaces the original sensitive attribute
A
i
with a fair dummy
A
~
i
sampled from the same data set (i.e., add cross-augmented dimensions to the modified training data), and subsequently updates the predictive model parameters
θ
f
(i.e., prior to training the first model with the modified training data) by using a loss function
l
Y
i
,
f
^
θ
f
X
i
and a penalty term, as shown in step 4 of Algorithm 1.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, WISNIEWSKI, and ROMANO before them, to include ROMANO’s replacement of sensitive attribute with a sampled fair dummy from the same data set in VERMA, LAM, and WISNIEWSKI’s bias removal method. One would have been motivated to make such a combination in order to learn predictive models that approximately satisfy the equalized odds notion of fairness (ROMANO [page 1, section 1 Introduction]) to achieve a more balanced performance across groups (ROMANO [page 2, section 1.1 A synthetic example]).
Regarding Claim 10:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. LAM further teaches:
wherein the one or more processors are further configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
VERMA in view of LAM and WISNIEWSKI is not relied upon for teaching, but ROMANO teaches: add cross-augmented data to replace the protected dimensions prior to training the first model with the modified training data, wherein the cross-augmented data includes values consistent with an original distribution of the protected dimensions. (ROMANO [page 1, section 1 Introduction] teaches: “Fairness constraints can often be articulated as conditional independence relations, and in this work we will focus on the equalized odds criterion [6], defined as
Y
^
A
|
Y
1
where the relationship above applies to test points; here,
Y
is the response variable,
A
is a sensitive attribute (e.g. gender),
X
is a vector of features that may also contain
A
, and
Y
^
=
f
^
X
is the prediction obtained with a fixed prediction rule
f
^
⋅
.” ROMANO [page 3, section 2.1 Regularization with fair dummies] teaches: “Our procedure starts by constructing a fair dummy sensitive attribute
A
~
i
for each training sample:
A
~
i
~
P
A
|
Y
A
i
|
Y
i
,
i
∈
I
t
r
a
i
n
,
where
P
A
|
Y
denotes the conditional distribution of
A
i
given
Y
i
. This sampling is straightforward; see (4) below. Importantly, we generate
A
~
i
without looking at
Y
^
i
so that we have the following property:
Y
^
i
A
~
i
|
Y
i
,
i
∈
I
t
r
a
i
n
,
2
Notice that the above is exactly the equalized odds relation in (1), with a crucial difference that the original sensitive attribute
A
i
is replaced by the artificial one
A
~
i
(i.e., wherein the cross-augmented data includes values consistent with an original distribution of the protected dimensions). […] we define
X
∈
R
I
t
r
a
i
n
×
p
,
A
~
∈
R
I
t
r
a
i
n
, and
Y
∈
R
I
t
r
a
i
n
, whose entries correspond to the features, sensitive attributes, fair dummies, and labels, respectively.” ROMANO [page 5, Algorithm 1] teaches step 2, which samples the fair dummies
A
~
i
~
P
A
|
Y
A
i
|
Y
i
and replaces the original sensitive attribute
A
i
with a fair dummy
A
~
i
sampled from the same data set (i.e., add cross-augmented data to replace the protected dimensions), and subsequently updates the predictive model parameters
θ
f
(i.e., prior to training the first model with the modified training data) by using a loss function
l
Y
i
,
f
^
θ
f
X
i
and a penalty term, as shown in step 4 of Algorithm 1.)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, WISNIEWSKI, and ROMANO before them, to include ROMANO’s replacement of sensitive attribute with a sampled fair dummy from the same data set in VERMA, LAM, and WISNIEWSKI’s bias removal method. One would have been motivated to make such a combination in order to learn predictive models that approximately satisfy the equalized odds notion of fairness (ROMANO [page 1, section 1 Introduction]) to achieve a more balanced performance across groups (ROMANO [page 2, section 1.1 A synthetic example]).
Regarding Claim 20:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 15 as outlined above Additionally, the claim recites similar limitations as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale.
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over VERMA in view of LAM and WISNIEWSKI as applied above to claim 8, and further in view of ALI (US 20240046012 A1), hereafter ALI.
Regarding Claim 11:
VERMA in view of LAM and WISNIEWSKI teaches the elements of claim 8 as outlined above. LAM further teaches:
wherein the one or more processors are further configured to: (LAM [0024] teaches: “According to various embodiments, the method 100 may be performed on one or more computing devices to train a machine learning model and then use the trained model to predict one or more unobserved outcome values in a way that reduces bias based on one or more protected attributes.” LAM [0047] teaches: “According to various embodiments, the method 400 may be implemented on any suitable computing device.” LAM [0108] teaches: “FIG. 9 illustrates one example of a computing device. According to various embodiments, a system 900 suitable for implementing embodiments described herein includes a processor 901, a memory module 903, a storage device 905, an interface 911, and a bus 915 (e.g., a PCI bus or other interconnection fabric.) System 900 may operate as a variety of devices such as computing device configured to perform data analysis, a cloud computing system configured to perform data analysis, or any other device or service described herein.”)
remove […] information, related to the protected dimensions, from the modified training data prior to training the first model with the modified training data. (LAM [0039] teaches: “For instance, features that have low predictive power but that are highly correlated with values corresponding with the protected attribute A 302 may be automatically removed (i.e., remove […] information, related to the protected dimensions).” LAM [0071] teaches: “If it is determined that the feature purely proxies for the selected attribute, then the selected feature is removed from the model and training data at 610 (i.e., from the modified training data).” LAM [0039] teaches: “In FIG. 3B, protected attribute pure proxies A′ 312 purely proxies for the protected attribute A 302 and are removed from the set of features used in both training (i.e., prior to training the first model with the modified training data) and inference.”)
VERMA in view of LAM and WISNIEWSKI is not relied upon for teaching, but ALI teaches: remove trend-based information, related to the […] (ALI [Abstract] teaches: “The memory device includes computer-executable instructions stored therein, which, when executed by the processor, cause the at least one processor to receive a plurality of historical data including one or more trends;” ALI [0016] teaches: “FIG. 8 illustrates an example process for pre-processing data to remove bias in accordance with at least one embodiment of the present disclosure.” ALI [0029] teaches: “This disclosure provides a synthetic data generator that receives a data set and trains a model to generate anonymized data to mimic the distribution of the original data set, without being tied to individual records.” ALI [0035] teaches: “For example, trends present in the fields of the real data 115 may be defined by marginal and conditional distributions of the real data 115 […].” ALI [0065] teaches: “FIG. 8 illustrates an example process 800 for pre-processing data to remove bias in accordance with at least one embodiment of the present disclosure.” ALI [0067] teaches: “In at least one embodiment, this anti-bias data pre-processing 420 is performed using a % K removal technique. In process 800, the data generation computer device 105 removes 805 features with high correlation to a protected attribute. In at least one embodiment, the high correlation is greater than or equal to 0.7 correlation with the protected attribute.” ALI [0071] teaches: “The data points of the synthetic data 310 along with the original data set 115 constitute an ideal dataset, as the labels no longer depend on protected attributes. Therefore, the model is trained on this overall dataset which represents an equitable world, thereby removing bias from the model.” ALI [0074] teaches: “For example, a ML module may receive training data comprising data associated with different trends and their corresponding classifications, generate a model which maps the trend data to the classification data, and recognize future trends and determine their corresponding categories.” Examiner’s note: ALI teaches a pre-processing method to remove bias in historical data including one or more trends by removing features with high correlation to a protected attribute (i.e., remove trend-based information, related to the protected dimensions).)
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of VERMA, LAM, WISNIEWSKI, and ALI before them, to include ALI’s historical training data including trends for removing bias to generate a data set that mimics the distribution of the original data set without being tied to individual records in VERMA, LAM, and WISNIEWSKI’s bias removal method. One would have been motivated to make such a combination in order to use training data comprising data associated with different trends and their corresponding classifications, generate a model which maps the trend data to the classification data, and recognize future trends and determine their corresponding categories (ALI [0074]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
SILBERMAN (US 20170330058 A1) relates to using machine learning to detect and reduce bias (including discrimination) in an automated decision making process.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Alvaro S Laham Bauzo whose telephone number is (571)272-5650. The examiner can normally be reached Mon-Fri 7:30 AM - 11:00 AM | 1:00 PM - 5:30 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.S.L./Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146