DETAILED ACTION
This action is in response to the initial filing of application no. 18/951,455 on 11/18/2024.
Claims 1 – 20 are still pending in this application, with claims 1,12 and 20 being independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claim 9 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 9 recites the following
receiving a plurality of contrastive questions for a plurality of contexts; receiving a prompt including at least two questions of the plurality of contrastive questions for a first context of the plurality of contexts, the at least two questions including a set of contradictory features associated with the first context; applying a large language model (LLM) on the prompt; generating a set of reasonings associated with the at least two questions, based on the application of the LLM on the prompt; generating a set of scores associated with the set of reasonings, based on the application of the LLM on the prompt; applying a statistical hypothesis testing model on the set of scores; determining whether the at least two questions including the set of contradictory features are statistically different for the first context, based on the application of the statistical hypothesis testing model; detecting a set of biases associated with the LLM, based on the determination that whether the at least two questions are statistically different; controlling rendering of first information including the set of biases associated with the LLM; determining a first median value and a first variance value corresponding to first scores of the set of scores, the first scores corresponding to a first question of the at least two questions for the first context; comparing the first median value with the second median value; and comparing the first variance value with the second variance value, wherein the detection of the set of biases associated with the LLM is further based on the comparison of the first median value with the second median value and the comparison of the first variance value with the second variance value.
After search and consideration of the prior art, it has been determined that the prior art fails to tech the following limitations: comparing the first variance value with the second variance value, wherein the detection of the set of biases associated with the LLM is further based on the comparison of the first median value with the second median value and the comparison of the first variance value with the second variance value. As discussed below, the combination of NIST and Murugan disclose determining median and variance values. However, the combination of NIST and Murugan fails to teach: comparing the first median value with the second median value; and comparing the first variance value with the second variance value, wherein the detection of the set of biases associated with the LLM is further based on the comparison of the first median value with the second median value and the comparison of the first variance value with the second variance value.
Claim 10 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 10 recites the following
receiving a plurality of contrastive questions for a plurality of contexts; receiving a prompt including at least two questions of the plurality of contrastive questions for a first context of the plurality of contexts, the at least two questions including a set of contradictory features associated with the first context; applying a large language model (LLM) on the prompt; generating a set of reasonings associated with the at least two questions, based on the application of the LLM on the prompt; generating a set of scores associated with the set of reasonings, based on the application of the LLM on the prompt; applying a statistical hypothesis testing model on the set of scores; determining whether the at least two questions including the set of contradictory features are statistically different for the first context, based on the application of the statistical hypothesis testing model; detecting a set of biases associated with the LLM, based on the determination that whether the at least two questions are statistically different; controlling rendering of first information including the set of biases associated with the LLM; determining a first median value and a first variance value corresponding to first scores of the set of scores, the first scores corresponding to a first question of the at least two questions for the first context; determining a first difference between the first median value and the second median value; and determining a second difference between the first variance value and the second variance value, wherein the detection of the set of biases associated with the LLM is further based on at least one of the first difference or the second difference.
After search and consideration of the prior art, it has been determined that the prior art fails to tech the following limitations: comparing the first variance value with the second variance value, wherein the detection of the set of biases associated with the LLM is further based on the comparison of the first median value with the second median value and the comparison of the first variance value with the second variance value. As discussed below, the combination of NIST and Murugan disclose determining median and variance values. However, the combination of NIST and Murugan fails to teach: determining a first difference between the first median value and the second median value; and determining a second difference between the first variance value and the second variance value, wherein the detection of the set of biases associated with the LLM is further based on at least one of the first difference or the second difference.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s)1 – 5, 12 – 15 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. (“Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models”) (“Kumar”) in view of Lohia et al. (US 2022/0237415) (“Lohia”) and further in view of Huang et al. (“TRUSTGPT: A Benchmark for Trustworthy and Responsible Large Language Models) (“Huang”).
For claim 1, Kumar discloses a method (Abstract) comprising: receiving a plurality of contrastive questions (a sentence comprising a gender related word that might results in a biased response and the same sentence comprising the gender counterpart (3. Gender Bias: Methods and Evaluation,3.1 Attacker LLM; Example-5A and Example-6A); receiving a prompt including at least two questions of the plurality of contrastive questions for a first context, the at least two questions including a set of contradictory features associated with the first context (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); applying a large language model (LLM) on the prompt (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of reasonings associated with the at least two questions (prompt response and CDA response), based on the application of the LLM on the prompt (Figure 2; 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of scores (bias scores for male and female responses) associated with the set of reasonings, based on the application of the LLM on the prompt (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); determining whether the at least two questions including the set of contradictory features are statistically different for the first context (Additionally, we calculate the difference in the LLM-as-a-Judge bias scores for male and female responses”, Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); detecting a set of biases associated with the LLM, based on the determination that whether the at least two questions are statistically different (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A).
Yet, Kumar fails to teach the following: receiving the questions for a plurality of contexts; applying a statistical hypothesis testing model on the set of scores (); determining the statistical difference based on the statistical hypotheses testing; controlling rendering of first information including the set of biases associated with the LLM; and executing the method using a processor.
However, Lohia discloses a method for providing priority-based, accuracy controlled individual fairness of unstructured text (Abstract), executed by a processor ([0004]), comprising the following: generating a plurality of contrastive (original and counterfactual) samples for a plurality of contexts (protected attributes) (age, gender, nationality) ([0016] [0017] [0019] [0020] [0029]); generating scores for contrastive samples based output of a machine learning model, wherein the scores indicate a relative level of bias between the one or more samples and the protected attribute ([0013] [0029] [0030); and controlling rendering of first information including the set of biases associated with a machine learning model (An unfairness quotient is calculated based on the difference in the scores. The original samples are ranked based according to the unfairness quotient. The original the machine learning model is re-trained using the counterfactual samples and original samples according to the ranking. Therefore, the rendering of information including the set of biases associated with the machine learning model is controlled by de-biasing the model by re-training the model. [0021 – 0023] [0029] [0030]).
Additionally, Huang discloses a method of evaluating a large language model in the area of bias (Abstract), comprising the following: a bias of a large language model is determined based on average score and statistical hypothesis testing (Figure 1; 1.Introduction Bias; 3.TRUSTGPT Benchmark, 3.1 Overall Design; 3.2 Models and Dataset, Model Selection; 3.3 Prompt Templates; 3.4 Metrics, 3.4.2 Bias; 4.2 Bias, pg. 3-6, 8 -9).
Therefore, it would have been obvious to one of ordinary of the art at the time of applicant’s filing to improve Kumar’s invention in the same way that Lohia’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]): the bias detection method (involving counterfactual data augmentation/testing) is further performed by a processor; the questions (text) are further received for a plurality of contexts (protected classes including gender and age) (Kumar, “In future work, we hope to explore these issues directly by expanding our work to other types of biases and protected classes.”, 5 Conclusion); the rendering of first information including the set of biases associated with a large language model is further controlled (by further de-biasing the model).
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kumar and Lohia in the same way that Huang’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract): a bias of a large language model is determined based on average score (Kumar, 3.3 Evaluation: LLM as a Judge) and statistical hypothesis testing (Kumar, 3.4 Evaluation: Other Metrics, Perspective API Metrics) (Huang, Figure1; 3.4.1 Toxicity and 4.2 Bias).
For claim 12, Kumar discloses a method (Abstract) comprising: receiving a plurality of contrastive questions (a sentence comprising a gender related word that might results in a biased response and the same sentence comprising the gender counterpart (3. Gender Bias: Methods and Evaluation,3.1 Attacker LLM; Example-5A and Example-6A); receiving a prompt including at least two questions of the plurality of contrastive questions for a first context, the at least two questions including a set of contradictory features associated with the first context (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); applying a large language model (LLM) on the prompt (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of reasonings associated with the at least two questions (prompt response and CDA response), based on the application of the LLM on the prompt (Figure 2; 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of scores (bias scores for male and female responses) associated with the set of reasonings, based on the application of the LLM on the prompt (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); determining whether the at least two questions including the set of contradictory features are statistically different for the first context (Additionally, we calculate the difference in the LLM-as-a-Judge bias scores for male and female responses”, Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); detecting a set of biases associated with the LLM, based on the determination that whether the at least two questions are statistically different (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A).
Yet, Kumar fails to teach the following: receiving the questions for a plurality of contexts; applying a statistical hypothesis testing model on the set of scores (); determining the statistical difference based on the statistical hypotheses testing; controlling rendering of first information including the set of biases associated with the LLM; and one or more non-transitory computer-readable storage medium configured to store instructions that, in response to being executed, causes an electronic device to perform the method.
However, Lohia discloses a method for providing priority-based, accuracy controlled individual fairness of unstructured text (Abstract), performed by a processor executing instructions stored on a non-transitory computer readable medium ([0004]), comprising the following: generating a plurality of contrastive (original and counterfactual) samples for a plurality of contexts (protected attributes) (age, gender, nationality) ([0016] [0017] [0019] [0020] [0029]); generating scores for contrastive samples based output of a machine learning model, wherein the scores indicate a relative level of bias between the one or more samples and the protected attribute ([0013] [0029] [0030); and controlling rendering of first information including the set of biases associated with a machine learning model (An unfairness quotient is calculated based on the difference in the scores. The original samples are ranked based according to the unfairness quotient. The original the machine learning model is re-trained using the counterfactual samples and original samples according to the ranking. Therefore, the rendering of information including the set of biases associated with the machine learning model is controlled by de-biasing the model by re-training the model. [0021 – 0023] [0029] [0030]).
Additionally, Huang discloses a method of evaluating a large language model in the area of bias (Abstract), comprising the following: a bias of a large language model is determined based on average score and statistical hypothesis testing (Figure 1; 1.Introduction Bias; 3.TRUSTGPT Benchmark, 3.1 Overall Design; 3.2 Models and Dataset, Model Selection; 3.3 Prompt Templates; 3.4 Metrics, 3.4.2 Bias; 4.2 Bias, pg. 3-6, 8 -9).
Therefore, it would have been obvious to one of ordinary of the art at the time of applicant’s filing to improve Kumar’s invention in the same way that Lohia’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]): the bias detection method (involving counterfactual data augmentation/testing) is further performed a processor executing instructions stored on a non-transitory computer readable medium; the questions (text) are further received for a plurality of contexts (protected classes including gender and age) (Kumar, “In future work, we hope to explore these issues directly by expanding our work to other types of biases and protected classes.”, 5 Conclusion); the rendering of first information including the set of biases associated with a large language model is further controlled (by further de-biasing the model).
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kumar and Lohia in the same way that Huang’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract): a bias of a large language model is determined based on average score (Kumar, 3.3 Evaluation: LLM as a Judge) and statistical hypothesis testing (Kumar, 3.4 Evaluation: Other Metrics, Perspective API Metrics) (Huang, Figure1; 3.4.1 Toxicity and 4.2 Bias).
For claim 2, Kumar further discloses, wherein the plurality of contrastive questions corresponds to at least one of: a curated dataset creation phase associated with the LLM, a problem formulation phase associated with the LLM, a data analysis phase associated with the LLM, or an evaluation phase associated with the LLM (Kumar, 3. Gender Bias: Methods and Evaluation, 3.1 Attacker LLM, 3.2 Target LLM, 3.3 Evaluation: LLM as a Judge, 3.4 Evaluation: Other Metrics).
For claims 3 and 13, Kumar and Lohia further disclose, wherein the first context corresponds to at least one of a gender stereotype context (Kumar, 3. Gender Bias: Methods and Evaluation) (Lohia, [0029] [0030]), a cultural context, an ethnicity context, a racial context, or a missing common-sense context.
For claims 4 and 14, Kumar and Lohia further disclose, wherein the set of biases corresponds to at least one of a gender stereotype bias (Kumar, 3. Gender Bias: Methods and Evaluation; C Sample Model Outputs with Evaluation Scores/Gaps, Example – 1A and Example-2A), a cultural bias, a confirmation or belief bias, an ethnicity bias, a racial bias, or a missing common-sense bias associated with the LLM.
For claims 5 and 15, Kumar and Huang further disclose: receiving a user input associated with a validation of the statistical difference (Kumar, 3.5 Human Evaluation) (Huang, , based on the determination that the at least two questions are statistically different (Kumar, 3.4 Evaluation: Other Metrics, Perspective API Metrics) (Huang, Figure1; 3.4.1 Toxicity and 4.2 Bias), wherein the detection of the set of biases associated with the LLM is further based on the received user input (Kumar, 3.5 Human Evaluation).
For claim 20, Kumar discloses a method (Abstract) comprising: receiving a plurality of contrastive questions (a sentence comprising a gender related word that might results in a biased response and the same sentence comprising the gender counterpart (3. Gender Bias: Methods and Evaluation,3.1 Attacker LLM; Example-5A and Example-6A); receiving a prompt including at least two questions of the plurality of contrastive questions for a first context, the at least two questions including a set of contradictory features associated with the first context (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); applying a large language model (LLM) on the prompt (Figure.2; Gender Bias: Methods and Evaluation,3.1 Attacker LLM, 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of reasonings associated with the at least two questions (prompt response and CDA response), based on the application of the LLM on the prompt (Figure 2; 3.2 Target LLM; 4. Results and Discussion; Example-5A and Example-6A); generating a set of scores (bias scores for male and female responses) associated with the set of reasonings, based on the application of the LLM on the prompt (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); determining whether the at least two questions including the set of contradictory features are statistically different for the first context (Additionally, we calculate the difference in the LLM-as-a-Judge bias scores for male and female responses”, Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A); detecting a set of biases associated with the LLM, based on the determination that whether the at least two questions are statistically different (Table1: Gender Bias Levels for LLM-as-a-Judge; Figure.2; Gender Bias: Methods and Evaluation, 3.3 Evaluation: LLM as a Judge; 4. Results and Discussion; Example-5A and Example-6A).
Yet, Kumar fails to teach the following: receiving the questions for a plurality of contexts; applying a statistical hypothesis testing model on the set of scores (); determining the statistical difference based on the statistical hypotheses testing; controlling rendering of first information including the set of biases associated with the LLM; and a memory configured to store instructions that are executed by a coupled processor to perform the method.
However, Lohia discloses a method for providing priority-based, accuracy controlled individual fairness of unstructured text (Abstract), performed by a processor executing instructions stored on a coupled memory([0004]), comprising the following: generating a plurality of contrastive (original and counterfactual) samples for a plurality of contexts (protected attributes) (age, gender, nationality) ([0016] [0017] [0019] [0020] [0029]); generating scores for contrastive samples based output of a machine learning model, wherein the scores indicate a relative level of bias between the one or more samples and the protected attribute ([0013] [0029] [0030); and controlling rendering of first information including the set of biases associated with a machine learning model (An unfairness quotient is calculated based on the difference in the scores. The original samples are ranked based according to the unfairness quotient. The original the machine learning model is re-trained using the counterfactual samples and original samples according to the ranking. Therefore, the rendering of information including the set of biases associated with the machine learning model is controlled by de-biasing the model by re-training the model. [0021 – 0023] [0029] [0030]).
Additionally, Huang discloses a method of evaluating a large language model in the area of bias (Abstract), comprising the following: a bias of a large language model is determined based on average score and statistical hypothesis testing (Figure 1; 1.Introduction Bias; 3.TRUSTGPT Benchmark, 3.1 Overall Design; 3.2 Models and Dataset, Model Selection; 3.3 Prompt Templates; 3.4 Metrics, 3.4.2 Bias; 4.2 Bias, pg. 3-6, 8 -9).
Therefore, it would have been obvious to one of ordinary of the art at the time of applicant’s filing to improve Kumar’s invention in the same way that Lohia’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]): the bias detection method (involving counterfactual data augmentation/testing) is further performed by a processor executing instructions stored on a coupled memory; the questions (text) are further received for a plurality of contexts (protected classes including gender and age) (Kumar, “In future work, we hope to explore these issues directly by expanding our work to other types of biases and protected classes.”, 5 Conclusion); the rendering of first information including the set of biases associated with a large language model is further controlled (by further de-biasing the model).
Additionally, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kumar and Lohia in the same way that Huang’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract): a bias of a large language model is determined based on average score (Kumar, 3.3 Evaluation: LLM as a Judge) and statistical hypothesis testing (Kumar, 3.4 Evaluation: Other Metrics, Perspective API Metrics) (Huang, Figure1; 3.4.1 Toxicity and 4.2 Bias).
Claim(s) 6 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. (“Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models”) (“Kumar”) in view of Lohia et al. (US 2022/0237415)(“Lohia”) and further in view of Huang et al. (“TRUSTGPT: A Benchmark for Trustworthy and Responsible Large Language Models) (“Huang”) and further in view of NIST (“Siegel Tukey Test”).
For claims 6 and 16, the combination of Kumar, Lohia and Huang fails to teach, wherein the statistical hypothesis testing model corresponds to a Siegal Tukey test model.
However, NIST discloses methods of analyzing real world data (Abstract), comprising the following: a Mann-Whitney test is performed as a step of a Siegel Turkey test (“Description: The Siegel-Tukey test is computed as follows: Combine the data from the two response variables and sort from smallest to largst. Assign rank 1 to the smallest observation, rank 2 to the largest observation, rank 3 to next largest observation, rank 4 to the second smallest observation. Continue this pattern until all values are ranked. If there are an odd number of total observations, then omit the observation that is the median of the combined observations. Perform a Mann-Whitney (rank sum) test on the ranked values.”).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kumar, Lohia and Huang in the same way that NIST’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract): the statistical hypothesis testing model (Huang, Mann-Whitney, 3.4 Metrics, 3.4.2 Bias; 4.2 Bias, 6.6.1 Mann-Whitney U test, pg. 3-6, 8 -9 and 16 – 17) is further performed as a step in a SIEGEL Tukey Test.
Claim(s) 7 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. (“Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models”) (“Kumar”) in view of Lohia et al. (US 2022/0237415)(“Lohia”) and further in view of Huang et al. (“TRUSTGPT: A Benchmark for Trustworthy and Responsible Large Language Models) (“Huang”), and further in view of NIST (“Siegel Tukey Test”) and further in view of Murugan (“Statistics: Range, Variance and Standard Deviation”).
For claims 7 and 17, the combination of Kumar, Lohia and Huang fails to teach the following: determining a median value and a variance value, associated with the set of scores, wherein the statistical hypothesis testing model is applied on the median value and the variance value.
However, NIST discloses methods of analyzing real world data (Abstract), comprising the following: a Mann-Whitney test is performed as a step of a Siegel Turkey test (“Description: The Siegel-Tukey test is computed as follows: Combine the data from the two response variables and sort from smallest to largst. Assign rank 1 to the smallest observation, rank 2 to the largest observation, rank 3 to next largest observation, rank 4 to the second smallest observation. Continue this pattern until all values are ranked. If there are an odd number of total observations, then omit the observation that is the median of the combined observations. Perform a Mann-Whitney (rank sum) test on the ranked values.”). Additionally, NIST discloses determining median and standard deviation values associated with a set of scores. The Siegel-Tukey test is applied on the median and standard deviation values (“ The Siegel-Tukey test is computed as follows: Combine the data from the two response variables and sort from smallest to largst. Assign rank 1 to the smallest observation, rank 2 to the largest observation, rank 3 to next largest observation, rank 4 to the second smallest observation. Continue this pattern until all values are ranked. If there are an odd number of total observations, then omit the observation that is the median of the combined observations. … Two Sample Two-Sided Siegel Tukey Test First Response Variable: Y1 Second Response Variable: Y2 H0: Sigma1 = Sigma2 Ha: Sigma1 not equal Sigma2 Summary Statistics: Number of Observations for Sample 1: 5 Mean for Sample 1: 16.05800 Median for Sample 1: 16.01000 Standard Deviation for Sample 1: 0.47007 Number of Observations for Sample: 5 Mean for Sample 2: 15.98400 Median for Sample 2: 15.98000 Standard Deviation for Sample 2: 0.09236”).
Additionally, Murugan discloses methods for providing statistics of real world data, wherein standard deviation is the square root of a variance (Variance and Standard Deviation).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s filing to improve the invention disclosed by the combination of Kumar, Lohia and Huang in the same way that NIST’s invention has been improved to achieve the following, predictable results for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract): the statistical hypothesis testing model (Huang, Mann-Whitney, 3.4 Metrics, 3.4.2 Bias; 4.2 Bias, 6.6.1 Mann-Whitney U test, pg. 3-6, 8 -9 and 16 – 17) is further performed as a step in a SIEGEL Tukey Test. Additionally, median and standard deviation values of the scores are determined, wherein the SIEGEL Tukey test is applied on these values.
Additionally, it would have been obvious to one of ordinary skill in the art the time of applicant’s filing to modify the combined teachings of Kumar, Lohia, Huang and NIST with Murugan’s teachings so that the standard deviation value is a variation value for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract) by testing and analyzing the results.
Claim(s) 8, 11, 18 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. (“Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models”) (“Kumar”) in view of Lohia et al. (US 2022/0237415)(“Lohia”) and further in view of Huang et al. (“TRUSTGPT: A Benchmark for Trustworthy and Responsible Large Language Models) (“Huang”) and further in view of Tamagnone et al. (“Leveraging Domain Knowledge for Inclusive and Bias-aware Humanitarian Response Entry Classification”) and further in view of Murugan (“Statistics: Range, Variance and Standard Deviation”).
For claims 8 and 18, the combination of Kumar, Lohia and Huang further discloses: determining a standard deviation value corresponding to first score of the set of scores (Kumar, 3.4 Evaluation Other Metrics )(Huang, 3.4.2 Bias and 6.6.4 Standard Deviation Calculation), the first scores corresponding to a first question of the at least two questions for the first context(Kumar, 3.4 Evaluation Other Metrics )(Huang, 3.4.2 Bias and 6.6.4 Standard Deviation Calculation); and determining a second standard deviation corresponding to second scores of the set of scores (Kumar, 3.4 Evaluation Other Metrics )(Huang, 3.4.2 Bias and 6.6.4 Standard Deviation Calculation), the second scores corresponding to a second question of the at least two questions for the first context (Kumar, 3.4 Evaluation Other Metrics )(Huang, 3.4.2 Bias and 6.6.4 Standard Deviation Calculation).
Yet, the combination of Kumar, Lohia and Huang fails to teach the following: determining a first median value and variance value corresponding to the first scores of the set of scores; and determining a second median value and variance value corresponding to the second scores of the set of scores.
However, Tamagnone discloses a method for detecting bias in a large language model (Abstract), wherein median values are determined for a sets of bias score generated by the large language model (“To measure the bias of the models, we use HUMSETBIAS to calculate the sensitivity of the trained models to the changes
in gender and countries keywords concerning the changes in the predicted probabilities of the labels. A non-biased model should not be sensitive to such changes so that the predicted probabilities should remain the same when the respective words of gender/country are swapped. To this end, our bias measurement is conducted on an extended version of the test sets, where each test set also includes the counterfactual samples of each data point (described in Section 4.2). Using this data, we first define the P-Shift metric as the shift in the predicted probability of the entry x concerning the tag t with the bias label m when replacing the label with n … Here, P(xm, t) is the prediction probability of x on the tag t in the form of bias label m, and P(xn, t) is the probability on the same tag t of a corresponding counterfactual sample of x
swapped to bias label n. The values are multiplied by 100 for improving readability. The total predicted probability shift from the attribute m to the attribute n for a tag t is therefore calculated as the median of the probability shifts.”, 4.1 HumBert, 4.2 HumSetBias, 5 Results and Analysis, 5.3 Bias Measurement and Mitigation).
Additionally, Murugan discloses methods for providing statistics of real world data, wherein standard deviation is the square root of a variance (Variance and Standard Deviation).
Therefore, it would have been obvious to one of ordinary skill in the art at the time of applicant’s invention to improve the combined teachings of Kumar, Lohia and Huang in the same way that Tamagnone’s invention has been improved to achieve the predictable results of further generating first and second median scores corresponding to the first and second questions, respectively for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract) by testing and analyzing the results using median values which could mitigate the effects of outlier values in the final statistic (Tamagnone, 5.3 Bias Measurement and Mitigation).
Additionally, it would have been obvious to one of ordinary skill in the art the time of applicant’s filing to modify the combined teachings of Kumar, Lohia, Huang and Tamagnone with Mutugan’s teachings so that the standard deviation value is a variation value for the purpose of controlling drops in accuracy while trying to achieve a sufficient level of fairness in the large language model (Lohia, [0013 – 0015]) (Huang, Abstract) by testing and analyzing the results.
For claims 11 and 19, Kumar and Huang further discloses: sorting the first scores and the second scores as a sorted list of scores (Kumar, 3 Gender Bias Evaluation, 3.2 Target LLM, 3.4 Evaluation Metrics; Example-1A and Example-2A) (Huang, 3.4 Metrics, 3.4.3 Bias; 4.2 Bias; 6.6 Metrics, Mann-Witney U Test, pg. 5, 6, 8 and 9); assigning alternate-extreme ranks to the sorted list of scores (Kumar, 3 Gender Bias Evaluation, 3.2 Target LLM, 3.4 Evaluation Metrics; Example-1A and Example-2A) (Huang, 3.4 Metrics, 3.4.3 Bias; 4.2 Bias; 6.6 Metrics, Mann-Witney U Test, pg. 5, 6, 8 and 9); and calculating a first rank sum for the first scores and a second rank sum for the second scores (Kumar, 3 Gender Bias Evaluation, 3.2 Target LLM, 3.4 Evaluation Metrics; Example-1A and Example-2A) (Huang, 3.4 Metrics, 3.4.3 Bias; 4.2 Bias; 6.6 Metrics, Mann-Witney U Test, pg. 5, 6, 8 and 9); based on the assignment of the alternate-extreme ranks, wherein the statistical hypothesis testing model is applied on the first rank sum and the second rank sum (Kumar, 3 Gender Bias Evaluation, 3.2 Target LLM, 3.4 Evaluation Metrics; Example-1A and Example-2A) (Huang, 3.4 Metrics, 3.4.3 Bias; 4.2 Bias; 6.6 Metrics, Mann-Witney U Test, pg. 5, 6, 8 and 9).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Paul et al. (US 2025/0356054) – discloses inventive concept of assessing gender fairness of LLMs using counterfactuals.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SONIA L GAY whose telephone number is (571)270-1951. The examiner can normally be reached Monday-Friday 9-5 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SONIA L GAY/Primary Examiner, Art Unit 2657