Prosecution Insights
Last updated: August 17, 2026
Application No. 18/599,284

NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM STORING INFORMATION PROCESSING PROGRAM, INFORMATION PROCESSING METHOD, AND INFORMATION PROCESSING APPARATUS

Non-Final OA §101§103§112
Filed
Mar 08, 2024
Priority
Sep 15, 2021 — continuation of PCTJP2021033991
Examiner
ADMASU, MAHLIET TASEW
Art Unit
Tech Center
Assignee
Fujitsu Limited
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
14 currently pending
Career history
12
Total Applications
across all art units

Statute-Specific Performance

§101
31.5%
-8.5% vs TC avg
§103
57.4%
+17.4% vs TC avg
§112
9.3%
-30.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This communication is in response to the Application No. 18/599,284 filed on March 08, 2024 in which Claims 1-13 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 1-13 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1, 12, and 13 recite the limitation "a relatively small number of data sets" in line 10. The term "relatively small" is a term of degree that is indefinite. The claim provides no reference point against which the number is measured, and it is unclear "relatively" to what the number is compared. Since independent claim 1 is rejected under 35 U.S.C. 112(b), claims 2 - 11 are also rejected under 35 U.S.C. 112(b) because they depend on claim 1 and inherit the same indefiniteness problems. Claim 2 recites the limitation "a relatively large number of data sets" in line 5. The term "relatively large" is a term of degree that is indefinite. The claim provides no reference point against which the number is measured, and it is unclear "relatively" to what the number is compared. Since independent claim 1 is rejected under 35 U.S.C. 112(b), claims 3, 6, 7, and 11 are also rejected under 35 U.S.C. 112(b) because they depend on claim 2 and inherit the same indefiniteness problems. Claim 3 recites the limitation "the target data" inline 5. There is insufficient antecedent basis for this limitation in the claim. The phrase "the target data" was not previously introduced. Claim 1, from which claim 3 depends, introduces "a target data set." It is unclear whether "the target data" refers to the previously recited "target data set" or to different subject matter. Claim 10 recites the limitation "the plurality of classifiers" in lines 4, 7, and 11. There is insufficient antecedent basis for this limitation in the claim. The phrase "the plurality of classifiers" was not previously introduced. Claim 1, from which claim 10 depends, recites only "a classifier" in the singular and does not introduce a plurality of classifiers. Claim 11 recites the limitation "the plurality of classifiers" in lines 4 and 7. There is insufficient antecedent basis for this limitation in the claim. The phrase "the plurality of classifiers" was not previously introduced. Claim 1, from which claim 11 depends, recites only "a classifier" in the singular and does not introduce a plurality of classifiers. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-13 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a non-transitory computer-readable recording medium type claim. Therefore, Claims 1-11 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. acquiring an index value, the index value indicating how many data sets have correct classification results obtained by classifying data sets for each of a plurality of attribute value patterns different from each other […] (mental process – acquiring an index value indicating how many data sets have correct classification results for each of a plurality of different attribute value patterns may be performed mentally or using pen and paper by a user observing the classification results associated with each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results for each attribute value pattern, and recording the resulting count as the index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern having a relatively small number of data sets with the correct classification results (mental process – identifying one or more first attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively small numbers of data sets with the correct classification results based on said analysis) in a case of classifying a target data set, determining whether at least any one of the identified one or more of first attribute value patterns matches the target data set (mental process - determining whether at least one identified first attribute value pattern matches the target data set may be performed mentally by a user observing/analyzing the attribute values of the target data set and the attribute values forming each identified first attribute value pattern, comparing the attribute values, and accordingly using judgment/evaluation to determine whether at least one of the identified first attribute value patterns matches the target data set based on said analysis) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 1 - 11. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. identifying, based on the acquired index values, one or more of second attribute value patterns among the plurality of attribute value patterns, each of the one or more of second attribute value patterns being an attribute value pattern having a relatively large number of data sets with the correct classification results (mental process - identifying one or more second attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively large numbers of data sets with the correct classification results based on said analysis) in the classifying of the target data set, determining whether at least any one of the identified one or more of second attribute value patterns matches the target data set (mental process -determining whether at least one identified second attribute value pattern matches the target data set may be performed mentally by a user observing/analyzing the attribute values of the target data set and the attribute values forming each identified second attribute value pattern, comparing the attribute values, and accordingly using judgment/evaluation to determine whether at least one of the identified second attribute value patterns matches the target data set based on said analysis) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 3 depends on. Step 2A Prong 2 & Step 2B: outputting first information indicating that a classification result of the target data set with the classifier is affirmed, when none of the identified one or more of first attribute value patterns matches the target data and at least any one of the identified one or more of second attribute value patterns matches the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 2. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Step 2A Prong 2 & Step 2B: outputting second information indicating that a classification result of the target data set with the classifier is denied, when at least any one of the identified one or more of first attribute value patterns matches the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 5 depends on. Step 2A Prong 2 & Step 2B: in the classifying of the target data set, when at least any one of the identified one or more of first attribute value patterns matches the target data set, outputting the first attribute value pattern matching the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 6 depends on. Step 2A Prong 2 & Step 2B: in the classifying of the target data set, when at least any one of the one or more of second attribute value patterns matches the target data set, outputting the second attribute value pattern matching the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 2. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 7 depends on. Step 2A Prong 2 & Step 2B: outputting first information indicating that a classification result of the target data set with the classifier is affirmed in association with the classification result, when none of the identified one or more of first attribute value patterns matches the target data set and at least any one of the identified one or more of second attribute value patterns matches the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 2. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: outputting second information indicating that a classification result of the target data set with the classifier is denied in association with the classification result, when at least any one of first attribute value patterns matches the target data set (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. acquiring the index value indicating how many data sets have correct classification results with the classifier in a case of classifying data sets for each of the plurality of attribute value patterns […](mental process - acquiring the index value indicating how many data sets have correct classification results with the classifier for each attribute value pattern may be performed mentally or using pen and paper by a user observing/analyzing the classification results associated with the data sets for each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results, and accordingly determining the index value based on said analysis) identifying, for each classifier of the plurality of classifiers, the one or more of first attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns being an attribute value pattern having a relatively small number of data sets with the correct classification results […] (mental process - identifying, for each classifier, one or more first attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns for each classifier, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively small numbers of data sets with the correct classification results based on said analysis) in the classifying of the target data set, selecting […a classifier…] with which at least any one of the one or more of first attribute value patterns matches the target data set (mental process - selecting a classifier with which at least one of the second attribute value patterns matches the target data set may be performed mentally or using pen and paper by a user observing/analyzing, for each classifier, the second attribute value patterns and the attribute values of the target data set, comparing the second attribute value patterns with the target data set, and accordingly using judgment/evaluation to select the classifier associated with at least one matching second attribute value pattern based on said analysis) Step 2A Prong 2 & Step 2B: […] with each of a plurality of classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier/classifiers without significantly more) […] obtained by the classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) […] and outputting among the plurality of classifiers, a classifier […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 10 depends on. in the classifying of data sets for each of the plurality of attribute value patterns […], acquiring, an index value indicating how many data sets have correct classification results […](mental process - acquiring an index value indicating how many data sets have correct classification results for each attribute value pattern may be performed mentally or using pen and paper by a user observing/analyzing the classification results associated with the data sets for each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results, and accordingly determining the index value based on said analysis) identifying, for each classifier of the plurality of classifiers, the one or more of second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of second attribute value patterns having a relatively large number of data sets with the correct classification results […] (mental process - identifying, for each classifier, one or more second attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns for each classifier, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively large numbers of data sets with the correct classification results based on said analysis) in the classifying of the target data set, selecting […a classifier…] with which at least any one of the one or more of second attribute value patterns matches the target data set (mental process - selecting a classifier with which at least one of the second attribute value patterns matches the target data set may be performed mentally or using pen and paper by a user observing/analyzing, for each classifier, the second attribute value patterns and the attribute values of the target data set, comparing the second attribute value patterns with the target data set, and accordingly using judgment/evaluation to select the classifier associated with at least one matching second attribute value pattern based on said analysis) Step 2A Prong 2 & Step 2B: […] by each classifier of the plurality of classifiers [...] obtained by the classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier/classifiers without significantly more) […] obtained by the classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) […] and outputting among the plurality of classifiers, a classifier […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 11 depends on. in the classifying of data sets for each of the plurality of attribute value patterns […], acquiring, an index value indicating how many data sets have correct classification results […](mental process - acquiring an index value indicating how many data sets have correct classification results for each attribute value pattern may be performed mentally or using pen and paper by a user observing/analyzing the classification results associated with the data sets for each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results, and accordingly determining the index value based on said analysis) identifying, for each classifier of the plurality of classifiers, the one or more of first attribute value patterns and the one or more second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns having a relatively small number of data sets with the correct classification results obtained by the classifier, each of the one or more of second attribute value patterns having a relatively large number of data sets with the correct classification results […] (mental process - identifying, for each classifier, the one or more first attribute value patterns and the one or more second attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns for each classifier, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively small numbers of data sets with the correct classification results and the attribute value patterns having relatively large numbers of data sets with the correct classification results based on said analysis) […] obtained by evaluating, based on the first attribute value pattern matching the target data set and the second attribute value pattern matching the target data set, a likelihood of a classification result of the target data set […](mental process - evaluating the likelihood of the classification result based on the first and second attribute value patterns matching the target data set may be performed mentally or using pen and paper by a user observing/analyzing the first attribute value pattern and the second attribute value pattern matching the target data set, considering the classification performance associated with the matching patterns, and accordingly using judgment/evaluation to evaluate the likelihood of the classification result based on said analysis) Step 2A Prong 2 & Step 2B: […] by each classifier of the plurality of classifiers [...] obtained by the classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier/classifiers without significantly more) […] obtained by the classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) in the classifying of the target data set, outputting a result […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] by each of the classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier/classifiers without significantly more) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 12: Step 1: Claim 12 is a method type claim. Therefore, Claim 12 falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. acquiring an index value, the index value indicating how many data sets have correct classification results obtained by classifying data sets for each of a plurality of attribute value patterns different from each other […] (mental process – acquiring an index value indicating how many data sets have correct classification results for each of a plurality of different attribute value patterns may be performed mentally or using pen and paper by a user observing the classification results associated with each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results for each attribute value pattern, and recording the resulting count as the index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern having a relatively small number of data sets with the correct classification results (mental process – identifying one or more first attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively small numbers of data sets with the correct classification results based on said analysis) in a case of classifying a target data set, determining whether at least any one of the identified one or more of first attribute value patterns matches the target data set (mental process - determining whether at least one identified first attribute value pattern matches the target data set may be performed mentally by a user observing/analyzing the attribute values of the target data set and the attribute values forming each identified first attribute value pattern, comparing the attribute values, and accordingly using judgment/evaluation to determine whether at least one of the identified first attribute value patterns matches the target data set based on said analysis) Step 2A Prong 2: This judicial exception is not integrated into a practical application. […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 12 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 13: Step 1: Claim 13 is an apparatus type claim. Therefore, Claims 13 falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. acquiring an index value, the index value indicating how many data sets have correct classification results obtained by classifying data sets for each of a plurality of attribute value patterns different from each other […] (mental process – acquiring an index value indicating how many data sets have correct classification results for each of a plurality of different attribute value patterns may be performed mentally or using pen and paper by a user observing the classification results associated with each attribute value pattern, comparing the classification results with the corresponding correct results, counting the number of correct classification results for each attribute value pattern, and recording the resulting count as the index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern having a relatively small number of data sets with the correct classification results (mental process – identifying one or more first attribute value patterns based on the acquired index values may be performed mentally or using pen and paper by a user observing/analyzing the numbers of correct classification results associated with the respective attribute value patterns, comparing the numbers of correct classification results, and accordingly using judgment/evaluation to identify the attribute value patterns having relatively small numbers of data sets with the correct classification results based on said analysis) in a case of classifying a target data set, determining whether at least any one of the identified one or more of first attribute value patterns matches the target data set (mental process - determining whether at least one identified first attribute value pattern matches the target data set may be performed mentally by a user observing/analyzing the attribute values of the target data set and the attribute values forming each identified first attribute value pattern, comparing the attribute values, and accordingly using judgment/evaluation to determine whether at least one of the identified first attribute value patterns matches the target data set based on said analysis) Step 2A Prong 2: This judicial exception is not integrated into a practical application. a memory; and a processor coupled to the memory […] (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. a memory; and a processor coupled to the memory […] (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with a classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying/using a classifier without significantly more) outputting a result of the determining (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 13 is rejected as being directed to an abstract idea without significantly more. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-13 are rejected under 35 U.S.C. 103 as being unpatentable over Chung et al. (hereafter Chung, a non-patent literature reference titled “Slice finder: Automated data slicing for model validation”) in view of Labisch et al. (hereinafter Labisch) (US 20210080924), and in further view of Chatterjee et al. (hereinafter Chatterjee) (US 20210049512). Regarding Claim 1, Chung teaches: acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs j Fj op vj where the Fj’s are distinct (e.g., country = DE ∧ gender = Male), and op can be one of =, . For numeric features, we can discretize their values (e.g., quantiles or equiheight bins) and generate ranges so that they are effectively categorical features (e.g., age = [20,30))” & Page 3 – Section II, “We also assume a classification loss function ψ(S,h) that returns a performance score for a set of examples by comparing h’s prediction h(x(i) F ) with the true label y(i). A common classification loss function is logarithmic loss (log loss)”, thus acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier is disclosed, because Chung teaches defining respective data slices using different combinations of feature values and obtaining a classification-performance score for the examples in each slice by comparing the classifier’s predictions with the true labels. Chung’s slices defined by conjunctions of feature value pairs correspond to the plurality of attribute value patterns different from each other, Chung’s examples within each slice correspond to the data sets, and Chung’s classification loss or performance score calculated for each slice corresponds to the acquired index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S, and the definition depends on the problem in hand”, & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S), which should deserve the user’s attention for deeper analysis”, thus identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] is disclosed, because Chung teaches comparing the classification loss value of each slice with that of its counterpart and identifying a problematic slice having a higher concentration of classification errors. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, and Chung’s identified problematic slices correspond to the one or more first attribute value patterns) […] the identified one or more of first attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h,” thus […] the identified one or more of first attribute value patterns […] is disclosed, because Chung teaches identifying one or more problematic slices having a higher concentration of classification errors. Chung’s identified problematic slices correspond to the identified one or more first attribute value patterns, and the common feature value pairs defining each problematic slice correspond to the attribute values forming each first attribute value pattern) Chung does not explicitly teach […] how many data sets have correct classification results […], […] a relatively small number of data sets with the correct classification results, in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Labisch teaches: […] how many data sets have correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] how many data sets have correct classification results […] is disclosed, because Labisch teaches detecting the classifications produced by each model, comparing those classifications with the underlying correct or true classifications, and generating a data set or histogram indicating the number of correct classifications. Labisch’s cases correspond to the data sets, the classifications matching the underlying correct or true classifications correspond to the correct classification results, and the number of correct classifications corresponds to how many data sets have correct classification results) […] having a relatively small number of data sets with the correct classification results (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] having a relatively small number of data sets with the correct classification results is disclosed, because Labisch teaches determining and comparing the number of cases in which classification results match the underlying correct or true classifications. Labisch’s cases correspond to the data sets, the matching classifications correspond to the correct classification results, and a comparatively lower number of correct classifications shown in the generated data set or histogram corresponds to a relatively small number of data sets with correct classification results) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chung with Labisch by using Labisch’s correct classification counting technique as the performance index for each of Chung’s data slices. Chung teaches classifying data sets for each of a plurality of attribute-value patterns different from each other, acquiring a performance value for each attribute-value pattern, and identifying one or more attribute-value patterns having relatively poor classification results. Labisch further teaches determining how many data sets have correct classification results by comparing the classification results obtained by a model with the underlying correct or true classifications and generating statistical information indicating the number of correct and incorrect classifications. Therefore, a POSITA would have been motivated to incorporate Labisch’s correct-classification counting technique into Chung’s process so that, for each of Chung’s plurality of attribute-value patterns different from each other, an index value indicating how many data sets have correct classification results could be acquired. The acquired index values could then be used to identify one or more first attribute-value patterns having a relatively small number of data sets with correct classification results, thereby providing a direct and reliable measure for identifying the attribute-value patterns on which the classifier performs poorly (Labisch, Par. [0031] “This can make it possible, in particular, to determine conditional confidences, which allow a particularly reliable assessment of the correctness of a classification”, & Par. [0032], “when performing the diagnosis method in accordance with the disclosed embodiments of the invention, the data set can be updated accordingly, i.e., the occurrence of incorrect or correct classifications incremented. Processing of an increasing volume of data makes it possible to improve the reliability of the determined confidence, therefore”) Chung combined with Labisch does not explicitly teach in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Chatterjee teaches: in a case of classifying a target data set, determining whether at least any one of […] matches the target data set (Chatterjee, Par. [0060], “After the classifier training is complete, a new observation record is received and a prediction 808 is provided for it. This post-training observation record 806 has a set of attribute values 807”, & Par. [0061], “the post-training observation record 806, the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, thus in a case of classifying a target data set, determining whether at least any one of […] matches the target data set is disclosed, because Chatterjee teaches receiving and classifying a new observation record, examining the identified rules to determine whether any rule is applicable, and determining whether the attribute values of the new observation record match the attribute predicates of at least one rule. Chatterjee’s new observation record corresponds to the target data set, and Chatterjee’s attribute predicate rules correspond to the identified attribute value patterns) and outputting a result of the determining (Chatterjee, Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”, & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a “no explanation is currently available” message and the updating of metrics or statistics regarding the effectiveness of the explainer”, thus and outputting a result of the determining is disclosed, because Chatterjee teaches outputting a matching rule as an explanation when the attribute values of the post-training observation record match the predicates of at least one rule, and outputting a “no explanation is currently available” message when the attribute values do not match any rule predicates. Chatterjee’s provided matching rule or no-explanation message corresponds to the output result of determining whether an identified attribute value pattern matches the target data set) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chung and Labisch with Chatterjee by applying Chatterjee’s attribute predicate matching process to the first attribute value patterns identified using Chung’s performance analysis and Labisch’s correct classification counts. Chung and Labisch collectively teach acquiring, for each of a plurality of attribute value patterns, an index value indicating how many data sets have correct classification results and identifying one or more first attribute value patterns having a relatively small number of data sets with correct classification results. Chatterjee further teaches, after classifying a new observation record, examining attribute predicate rules to determine whether any rule matches the attribute values of the new observation record and outputting either the matching rule or an indication that no rule applies. Therefore, a POSITA would have been motivated to compare a target data set with the identified first attribute value patterns and output the result of the comparison so that the classification result of the target data set could be evaluated based on whether the target data set matches a previously identified attribute value pattern associated with relatively few correct classification results. This would allow the system to inform the user whether the target data set falls within a pattern for which the classifier has demonstrated relatively poor classification performance (Chatterjee, Par. [0061], “The explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”) Regarding Claim 2, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: identifying, based on the acquired index values, one or more of second attribute value patterns among the plurality of attribute value patterns, each of the one or more of second attribute value patterns being an attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ' ),” thus identifying, based on the acquired index values, one or more of second attribute value patterns among the plurality of attribute value patterns, each of the one or more of second attribute value patterns being an attribute value pattern […] is disclosed, because Chung teaches comparing the classification loss values of a slice and its counterpart and distinguishing the counterpart having a relatively lower concentration of classification errors. Chung’s classification loss values correspond to the acquired index values, Chung’s slices and counterpart slices defined by feature value pairs correspond to the plurality of attribute value patterns, and Chung’s relatively better performing counterpart slices correspond to the one or more second attribute value pattern) [...] the identified one or more of second attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ' ),” thus […] the identified one or more of second attribute value patterns […] is disclosed, because Chung teaches identifying and distinguishing counterpart slices having relatively lower error concentrations than the problematic slices. Chung’s counterpart slices correspond to the identified one or more second attribute value patterns, and the common feature value pairs defining the counterpart slices correspond to the attribute values forming each second attribute value pattern) Labisch teaches: […] having a relatively large number of data sets with the correct classification results (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] having a relatively large number of data sets with the correct classification results is disclosed, because Labisch teaches determining and comparing the number of cases in which the classification results match the underlying correct or true classifications. Labisch’s cases correspond to the data sets, the matching classifications correspond to the correct classification results, and a comparatively higher number of correct classifications shown in the generated data set or histogram corresponds to a relatively large number of data sets with correct classification results) Chatterjee teaches: in the classifying of the target data set, determining whether at least any one of […] matches the target data set (Chatterjee, Par. [0060], “After the classifier training is complete, a new observation record is received and a prediction 808 is provided for it. This post training observation record 806 has a set of attribute values 807,” & Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction” thus in the classifying of the target data set, determining whether at least any one of […] matches the target data set is disclosed, because Chatterjee teaches receiving and classifying a new observation record, examining the identified rules, and determining whether the attribute values of the new observation record match the attribute predicates of at least one rule. Chatterjee’s new observation record corresponds to the target data set, and Chatterjee’s attribute predicate rules correspond to the identified attribute value patterns) Regarding Claim 3, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 2 as cited above and Chung further teaches: […] the identified one or more of first attribute value patterns […] the identified one or more of second attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ' ),” thus […] the identified one or more of first attribute value patterns […] the identified one or more of second attribute value patterns […] are disclosed, because Chung teaches identifying problematic slices having relatively higher concentrations of classification errors and distinguishing the counterpart slices having relatively lower concentrations of classification errors. Chung’s identified problematic slices correspond to the identified one or more of first attribute value patterns, and Chung’s relatively better-performing counterpart slices correspond to the identified one or more of second attribute value patterns) Chatterjee teaches: outputting first information indicating that a classification result of the target data set with the classifier is affirmed, when none of […] matches the target data and at least any one of […] matches the target data set (Chatterjee, Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”, & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a “no explanation is currently available” message and the updating of metrics or statistics regarding the effectiveness of the explainer”, thus outputting first information indicating that a classification result of the target data set with the classifier is affirmed, when none of […] matches the target data and at least any one of […] matches the target data set is disclosed, because Chatterjee teaches examining the identified rules to determine whether they apply to the target observation, determining when the target observation does not match specified rule predicates, and providing a matching rule as an explanation when the target observation matches at least one rule and the classifier’s prediction matches the implication of that rule. Chatterjee’s determination that no applicable rule associated with […] matches corresponds to none of […] matching the target data, Chatterjee’s determination that at least one applicable rule matches corresponds to at least any one of […] matching the target data set, and Chatterjee’s matching rule provided as an explanation for the prediction corresponds to the first information indicating that the classification result is affirmed) Regarding Claim 4, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: […] the identified one or more of first attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h,” thus […] the identified one or more of first attribute value patterns […] is disclosed, because Chung teaches identifying one or more problematic slices having a relatively higher concentration of classification errors. Chung’s identified problematic slices correspond to the identified one or more of first attribute value patterns) Chatterjee teaches: outputting second information indicating that a classification result of the target data set with the classifier is denied, when at least any one of […] matches the target data set (Chatterjee, Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus outputting second information indicating that a classification result of the target data set with the classifier is denied, when at least any one of […] matches the target data set is disclosed by the combined teachings, because Chatterjee teaches determining that the target data set matches at least one identified attribute predicate rule and outputting the matching rule as information concerning the classification result. When Chatterjee’s matching process is applied to Chung’s identified problematic slices, the output of a matching problematic slice indicates that the target data set falls within an attribute value pattern having relatively poor classification performance and therefore corresponds to second information indicating that the classification result is denied) Regarding Claim 5, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: […] the identified one or more of first attribute value patterns [...] the first attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ' ),” thus […] the identified one or more of first attribute value patterns [...] the first attribute value pattern […] are disclosed, because Chung teaches identifying one or more problematic slices having relatively higher concentrations of classification errors. Chung’s identified problematic slices correspond to the identified one or more of first attribute value patterns, and an individual identified problematic slice corresponds to the first attribute value pattern) Chatterjee teaches: in the classifying of the target data set, when at least any one of […] matches the target data set, outputting […] matching the target data set (Chatterjee, Par. [0060], “After the classifier training is complete, a new observation record is received and a prediction 808 is provided for it. This post-training observation record 806 has a set of attribute values 807,” & Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus in the classifying of the target data set, when at least any one of […] matches the target data set, outputting […] matching the target data set is disclosed, because Chatterjee teaches classifying a new observation record, determining whether its attribute values match the attribute predicates of at least one identified rule, and outputting the matching rule as an explanation for the prediction. Chatterjee’s new observation record corresponds to the target data set, determining that at least one rule applies corresponds to determining that at least any one of […] matches the target data set, and Chatterjee’s output matching rule corresponds to outputting […] matching the target data set) Regarding Claim 6, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 2 as cited above and Chung further teaches: […] the one or more of second attribute value patterns […] the second attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ' ),” thus […] the one or more of second attribute value patterns […] the second attribute value pattern […] are disclosed, because Chung teaches distinguishing one or more counterpart slices having relatively lower concentrations of classification errors than the problematic slices. Chung’s counterpart slices correspond to the one or more of second attribute value patterns, and an individual counterpart slice corresponds to the second attribute value pattern) Chatterjee teaches: in the classifying of the target data set, when at least any one of […] matches the target data set, outputting […] matching the target data set (Chatterjee, Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus in the classifying of the target data set, when at least any one of […] matches the target data set, outputting […] matching the target data set is disclosed, because Chatterjee teaches determining that the attribute values of the target observation match the attribute predicates of at least one rule and providing that matching rule to the client as an explanation for the prediction. Chatterjee’s determination that the attribute values match at least one rule corresponds to determining that at least any one of […] matches the target data set, and the rule provided to the client corresponds to outputting […] matching the target data set) Regarding Claim 7, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 2 as cited above and Chung further teaches: […] the identified one or more of first attribute value patterns […] the identified one or more of second attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h ,” thus […] the identified one or more of first attribute value patterns […] the identified one or more of second attribute value patterns […] are disclosed, because Chung teaches identifying problematic slices having relatively higher concentrations of classification errors and distinguishing their counterpart slices having relatively lower concentrations of classification errors. Chung’s identified problematic slices correspond to the identified one or more of first attribute value patterns, and Chung’s relatively better performing counterpart slices correspond to the identified one or more of second attribute value patterns) Chatterjee teaches: outputting first information indicating that a classification result of the target data set with the classifier is affirmed in association with the classification result, when none of […] matches the target data set and at least any one of […] matches the target data set (Chatterjee, Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a ‘no explanation is currently available’ message,” thus outputting first information indicating that a classification result of the target data set with the classifier is affirmed in association with the classification result, when none of […] matches the target data set and at least any one of […] matches the target data set is disclosed by the combined teachings, because Chatterjee teaches determining that the target observation does not match one or more rule predicates, determining that it matches at least one other rule whose implication agrees with the classifier’s prediction, and providing the matching rule as an explanation for that prediction. When Chatterjee’s matching process is applied to Chung’s first and second attribute value patterns, the absence of a match with the first attribute value patterns corresponds to none of […] matching the target data set, the match with a second attribute value pattern corresponds to at least any one of […] matching the target data set, and the matching rule provided as an explanation for the prediction corresponds to the first information affirming the classification result in association with that classification result) Regarding Claim 8, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: […] with the classifier […] one of first attribute value patterns […] (Chung, Page 3 – Section II, “We also assume a classification loss function ψ(S,h) that returns a performance score for a set of examples by compar ing h’s prediction h(x(i) F ) with the true label y(i)”, & Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S, and the definition depends on the problem in hand. For instance, in the most general case where user wants to validate if the model under performs on any data slices, we define the counterpart as the complement of S (S = D −S) and consider the difference ψ(S,h) − ψ(S,h) assuming ψ is a loss function, such as a log loss. (The definition of counterpart can change in other scenarios as we explain later.) This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S), which should deserve the user’s attention for deeper analysis”, thus […] with the classifier […] one of first attribute value patterns […] is disclosed, because Chung teaches evaluating the predictions generated by classifier h for examples contained in respective data slices and identifying a particular problematic slice S having a higher concentration of classification errors than its counterpart. Chung’s classifier h corresponds to the classifier, and Chung’s particular problematic slice S , defined by common feature value pairs and identified based on the classification loss of classifier h , corresponds to one of first attribute value patterns) Chatterjee teaches: outputting second information indicating that a classification result of the target data set […] is denied in association with the classification result, when at least any […] matches the target data set (Chatterjee, Abstract, “A set of explanatory rules is mined from the transformed data set, with each rule indicating a relationship between the prediction and one or more features corresponding to the training records. From among the rule set, a particular matching rule is selected to provide an easy-to-understand explanation for a prediction made by the classifier for an observation record which is not part of the training set,” & Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus outputting information in association with a classification result of the target data set, when at least any […] matches the target data set is disclosed, because Chatterjee teaches determining whether a rule applies to a target observation based on whether the observation’s attribute values match the rule’s attribute predicates and, upon identifying a matching rule, providing that rule to the client as an explanation associated with the classifier’s prediction. Chatterjee’s determination that at least one rule applies corresponds to determining that at least any […] matches the target data set, and the matching rule provided as an explanation for the prediction corresponds to outputting information in association with the classification result) Regarding Claim 9, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: the acquiring includes acquiring the index value indicating […] with the classifier in a case of classifying data sets for each of the plurality of attribute value patterns […] (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs,” & Page 3 – Section II, “We also assume a classification loss function ψ ( S , h ) that returns a performance score for a set of examples by comparing h ’s prediction h ( x F i ) with the true label y i ,” thus the acquiring includes acquiring the index value indicating […] with the classifier in a case of classifying data sets for each of the plurality of attribute value patterns […] is disclosed, because Chung teaches acquiring a performance score based on the predictions generated by classifier h for the examples contained in each respective slice. Chung’s performance score corresponds to the index value, classifier h corresponds to the classifier, the examples correspond to the data sets, and the slices defined by common feature value pairs correspond to the plurality of attribute value patterns) the identifying includes […] the one or more of first attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns being an attribute value pattern […] obtained by the classifier (Chung, Page 3 – Section II, “We also assume a classification loss function ψ(S,h) that returns a performance score for a set of examples by compar ing h’s prediction h(x(i) F ) with the true label y(i)”, & Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S, and the definition depends on the problem in hand. For instance, in the most general case where user wants to validate if the model under performs on any data slices, we define the counterpart as the complement of S (S = D −S) and consider the difference ψ(S,h) − ψ(S,h) assuming ψ is a loss function, such as a log loss. (The definition of counterpart can change in other scenarios as we explain later.) This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S), which should deserve the user’s attention for deeper analysis”, thus the identifying includes […] the one or more of first attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns being an attribute value pattern […] obtained by the classifier is disclosed, because Chung teaches acquiring a classification performance score for each slice by comparing the predictions generated by classifier h with the corresponding true labels and identifying a problematic slice S having a higher concentration of classification errors based on the acquired classification loss values. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, Chung’s identified problematic slices correspond to the one or more of first attribute value patterns, and Chung’s predictions generated by classifier h correspond to classification results obtained by the classifier) […] the one or more of first attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart,” & “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S ),” thus […] the one or more of first attribute value patterns […] is disclosed, because Chung teaches identifying one or more problematic slices having relatively higher concentrations of classification errors. Chung’s identified problematic slices, which are defined by common feature value pairs, correspond to the one or more of first attribute value patterns) Labisch teaches: […] how many data sets have correct classification results […] with each of a plurality of classifiers (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] how many data sets have correct classification results […] with each of a plurality of classifiers is disclosed, because Labisch teaches determining, for each of multiple classification models, the number of cases in which the classification generated by the model matches the underlying correct or true classification. Labisch’s cases correspond to the data sets, the classifications matching the correct or true classifications correspond to the correct classification results, and Labisch’s multiple models correspond to the plurality of classifiers) […] identifying, for each classifier of the plurality of classifiers, […] having a relatively small number of data sets with the correct classification results [...] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] identifying, for each classifier of the plurality of classifiers, […] having a relatively small number of data sets with the correct classification results [...] is disclosed, because Labisch teaches determining and comparing, for each of multiple classification models, the number of cases in which the classification produced by the respective model matches the underlying correct or true classification. Labisch’s multiple models correspond to the plurality of classifiers, each respective model corresponds to each classifier, Labisch’s cases correspond to the data sets, and a comparatively lower number of classifications matching the underlying correct or true classification corresponds to a relatively small number of data sets with the correct classification results) […] selecting and outputting among the plurality of classifiers, a classifier […] (Labisch, Claim 1, “classifying the status of the process-engineering plant aided by a plurality of models,” & “outputting diagnosis information which is based on the classifications of the status of the process-engineering plant and a confidence allocated thereto,” & Claim 10, “the overall classification is determined on the basis of the classifications, which result from the at least two models, and the confidences allocated thereto,” thus […] selecting and outputting among the plurality of classifiers, a classifier […] is disclosed, because Labisch teaches generating classifications using a plurality of models, selecting an overall classification based on the classifications produced by the respective models and their associated confidences, and outputting diagnosis information based on the selected classification. Labisch’s plurality of models corresponds to the plurality of classifiers, and the model whose classification is selected for the overall classification corresponds to the selected and output classifier) Chatterjee teaches: in the classifying of the target data set, […] with which none of […] matches the target data set (Chatterjee, Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a ‘no explanation is currently available’ message and the updating of metrics or statistics regarding the effectiveness of the explainer,” thus in the classifying of the target data set, […] with which none of […] matches the target data set is disclosed, because Chatterjee teaches examining the rules during classification of a post training observation record and determining that none of the generated rule predicates matches the attribute values of the observation record. Chatterjee’s post training observation record corresponds to the target data set, and Chatterjee’s determination that none of the rule predicates matches the attribute values corresponds to determining that none of […] matches the target data set) Regarding Claim 10, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: the acquiring includes in the classifying of data sets for each of the plurality of attribute value patterns […] acquiring, an index value indicating […] obtained by the classifier (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs,” & Page 3 – Section II, “We also assume a classification loss function ψ ( S , h ) that returns a performance score for a set of examples by comparing h ’s prediction h ( x F i ) with the true label y i ,” thus the acquiring includes in the classifying of data sets for each of the plurality of attribute value patterns […] acquiring, an index value indicating […] obtained by the classifier is disclosed, because Chung teaches defining respective data slices using different combinations of feature values and acquiring a classification performance score for the examples in each slice based on predictions generated by classifier h . Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, Chung’s examples correspond to the data sets, Chung’s classification performance score corresponds to the acquired index value, and Chung’s predictions generated by classifier h correspond to classification results obtained by the classifier) the identifying includes identifying, […] the one or more of second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of second attribute value patterns […] obtained by the classifier (Chung, Page 3 – Section II, “We also assume a classification loss function ψ ( S , h ) that returns a performance score for a set of examples by comparing h ’s prediction h ( x F i ) with the true label y i ,” & Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & “This effectively allows us to identify S with a higher error concentration for h ,” thus the identifying includes identifying, […] the one or more of second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of second attribute value patterns […] obtained by the classifier is disclosed, because Chung teaches comparing the classification loss values acquired for respective slices and distinguishing counterpart slices having relatively lower concentrations of classification errors than the identified problematic slices. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, Chung’s counterpart slices having relatively lower error concentrations correspond to the one or more of second attribute value patterns, and Chung’s predictions generated by classifier h correspond to classification results obtained by the classifier) […] the one or more of second attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & “This effectively allows us to identify S with a higher error concentration for h ,” thus […] the one or more of second attribute value patterns […] is disclosed, because Chung teaches distinguishing counterpart slices having relatively lower concentrations of classification errors than the problematic slices. Chung’s counterpart slices defined by common feature value pairs correspond to the one or more of second attribute value patterns) Labisch teaches: […] by each classifier of the plurality of classifiers, […] how many data sets have correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] by each classifier of the plurality of classifiers, […] how many data sets have correct classification results […] is disclosed, because Labisch teaches determining, for each of multiple classification models, the number of cases in which the classification generated by the respective model matches the underlying correct or true classification. Labisch’s multiple models correspond to the plurality of classifiers, each respective model corresponds to each classifier, Labisch’s cases correspond to the data sets, and the number of classifications matching the correct or true classifications corresponds to how many data sets have correct classification results) […] for each classifier of the plurality of classifiers, […] having a relatively large number of data sets with the correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] for each classifier of the plurality of classifiers, […] having a relatively large number of data sets with the correct classification results […] is disclosed, because Labisch teaches determining and comparing, for each of multiple classification models, the number of cases in which the classification produced by the respective model matches the underlying correct or true classification. Labisch’s multiple models correspond to the plurality of classifiers, each respective model corresponds to each classifier, Labisch’s cases correspond to the data sets, and a comparatively higher number of classifications matching the underlying correct or true classification corresponds to a relatively large number of data sets with the correct classification results) […] selecting and outputting among the plurality of classifiers, a classifier […] (Labisch, Claim 1, “classifying the status of the process-engineering plant aided by a plurality of models,” & “outputting diagnosis information which is based on the classifications of the status of the process-engineering plant and a confidence allocated thereto,” & Claim 10, “the overall classification is determined on the basis of the classifications, which result from the at least two models, and the confidences allocated thereto,” thus selecting and outputting among the plurality of classifiers, a classifier […] is disclosed, because Labisch teaches producing classifications using a plurality of models, selecting an overall classification based on the classifications generated by the respective models and their associated confidences, and outputting diagnosis information based on the selected classification. Labisch’s plurality of models corresponds to the plurality of classifiers, and the model associated with the selected overall classification corresponds to the selected and output classifier) Chatterjee teaches: in the classifying of the target data set, […] with which at least any one of […] matches the target data set (Chatterjee, Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus in the classifying of the target data set, […] with which at least any one of […] matches the target data set is disclosed, because Chatterjee teaches examining the rules associated with a classifier during classification of a target observation and determining that the attribute values of the target observation match the attribute predicates of at least one rule. Chatterjee’s post training observation record corresponds to the target data set, and determining that at least one rule associated with the classifier applies corresponds to determining that at least any one of […] matches the target data set with that classifier) Regarding Claim 11, Chung and Labisch combined with Chatterjee teaches all the limitations of claim 1 as cited above and Chung further teaches: the acquiring includes in the classifying of data sets for each of the plurality of attribute value patterns […] acquiring, an index value indicating […] obtained by the classifier (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs,” & Page 3 – Section II, “We also assume a classification loss function ψ ( S , h ) that returns a performance score for a set of examples by comparing h ’s prediction h ( x F i ) with the true label y i ,” thus the acquiring includes in the classifying of data sets for each of the plurality of attribute value patterns […] acquiring, an index value indicating […] obtained by the classifier is disclosed, because Chung teaches defining respective data slices using different combinations of feature values and acquiring a performance score for the examples in each slice based on predictions generated by classifier h . Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, Chung’s examples correspond to the data sets, Chung’s performance score corresponds to the acquired index value, and Chung’s predictions generated by classifier h correspond to classification results obtained by the classifier) the identifying includes identifying, […], the one or more of first attribute value patterns and the one or more second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns […] obtained by the classifier, each of the one or more of second attribute value patterns […] obtained by the classifier (Chung, Page 3 – Section II, “We also assume a classification loss function ψ ( S , h ) that returns a performance score for a set of examples by comparing h ’s prediction h ( x F i ) with the true label y i ,” & Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h ,” thus the identifying includes identifying, […] the one or more of first attribute value patterns and the one or more second attribute value patterns among the plurality of attribute value patterns based on the acquired index values, each of the one or more of first attribute value patterns […] obtained by the classifier, each of the one or more of second attribute value patterns […] obtained by the classifier is disclosed, because Chung teaches comparing classification loss values acquired for respective slices and distinguishing problematic slices having higher concentrations of classification errors from their counterpart slices having relatively lower concentrations of classification errors. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, Chung’s problematic slices correspond to the one or more of first attribute value patterns, Chung’s counterpart slices correspond to the one or more second attribute value patterns, and Chung’s predictions generated by classifier h correspond to classification results obtained by the classifier) […] based on the first attribute value pattern […] and the second attribute value pattern, […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S ,” & “This effectively allows us to identify S with a higher error concentration for h ,” thus […] based on the first attribute value pattern […] and the second attribute value pattern, […] is disclosed, because Chung teaches evaluating a problematic slice having a higher concentration of classification errors relative to its counterpart slice having a lower concentration of classification errors. Chung’s problematic slice corresponds to the first attribute value pattern, and Chung’s counterpart slice corresponds to the second attribute value pattern) Labisch teaches: […] by each classifier of the plurality of classifiers, […] how many data sets have correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] by each classifier of the plurality of classifiers, […] how many data sets have correct classification results […] is disclosed, because Labisch teaches determining, for each of multiple classification models, the number of cases in which the classification generated by the respective model matches the underlying correct or true classification. Labisch’s multiple models correspond to the plurality of classifiers, each respective model corresponds to each classifier, Labisch’s cases correspond to the data sets, and the number of classifications matching the underlying correct or true classification corresponds to how many data sets have correct classification results) […] for each classifier of the plurality of classifiers […] having a relatively small number of data sets with the correct classification results […] having a relatively large number of data sets with the correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges,” thus […] for each classifier of the plurality of classifiers […] having a relatively small number of data sets with the correct classification results […] having a relatively large number of data sets with the correct classification results […] is disclosed, because Labisch teaches determining and comparing, for each of multiple classification models, the number of cases in which the classification produced by the respective model matches the underlying correct or true classification and generating statistical information showing the number of correct classifications for each model. Labisch’s multiple models correspond to the plurality of classifiers, each respective model corresponds to each classifier, Labisch’s cases correspond to the data sets, a comparatively lower number of classifications matching the correct or true classification corresponds to a relatively small number of data sets with the correct classification results, and a comparatively higher number of classifications matching the correct or true classification corresponds to a relatively large number of data sets with the correct classification results) […] outputting a result obtained by evaluating, […] a likelihood of a classification result of the target data set by each of the classifiers (Labisch, Claim 1, “classifying the status of the process-engineering plant aided by a plurality of models,” & “outputting diagnosis information which is based on the classifications of the status of the process-engineering plant and a confidence allocated thereto,” & Claim 10, “the overall classification is determined on the basis of the classifications, which result from the at least two models, and the confidences allocated thereto,” thus […] outputting a result obtained by evaluating, […] a likelihood of a classification result of the target data set by each of the classifiers is disclosed, because Labisch teaches evaluating the classifications generated by multiple models based on a confidence allocated to each classification, determining an overall classification based on the respective classifications and confidences, and outputting diagnosis information based thereon. Labisch’s plurality of models corresponds to the classifiers, the confidence allocated to each model’s classification corresponds to the likelihood of the classification result by each classifier, and the overall classification and diagnosis information correspond to the output result) Chatterjee teaches: in the classifying of the target data set, […] matching the target data set […] matching the target data set […] (Chatterjee, Par. [0061], “the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not,” & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction,” thus in the classifying of the target data set, […] matching the target data set […] matching the target data set […] is disclosed, because Chatterjee teaches examining rules during classification of a target observation and determining whether the observation’s attribute values match the attribute predicates of the applicable rules. Chatterjee’s observation record corresponds to the target data set, and the attribute predicate rules matching the observation’s attribute values correspond to the respective patterns matching the target data set) Regarding Claim 12, Chung teaches: acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs j Fj op vj where the Fj’s are distinct (e.g., country = DE ∧ gender = Male), and op can be one of =, . For numeric features, we can discretize their values (e.g., quantiles or equiheight bins) and generate ranges so that they are effectively categorical features (e.g., age = [20,30))” & Page 3 – Section II, “We also assume a classification loss function ψ(S,h) that returns a performance score for a set of examples by comparing h’s prediction h(x(i) F ) with the true label y(i). A common classification loss function is logarithmic loss (log loss)”, thus acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier is disclosed, because Chung teaches defining respective data slices using different combinations of feature values and obtaining a classification-performance score for the examples in each slice by comparing the classifier’s predictions with the true labels. Chung’s slices defined by conjunctions of feature value pairs correspond to the plurality of attribute value patterns different from each other, Chung’s examples within each slice correspond to the data sets, and Chung’s classification loss or performance score calculated for each slice corresponds to the acquired index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S, and the definition depends on the problem in hand”, & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S), which should deserve the user’s attention for deeper analysis”, thus identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] is disclosed, because Chung teaches comparing the classification loss value of each slice with that of its counterpart and identifying a problematic slice having a higher concentration of classification errors. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, and Chung’s identified problematic slices correspond to the one or more first attribute value patterns) […] the identified one or more of first attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h,” thus […] the identified one or more of first attribute value patterns […] is disclosed, because Chung teaches identifying one or more problematic slices having a higher concentration of classification errors. Chung’s identified problematic slices correspond to the identified one or more first attribute value patterns, and the common feature value pairs defining each problematic slice correspond to the attribute values forming each first attribute value pattern) Chung does not explicitly teach […] how many data sets have correct classification results […], […] a relatively small number of data sets with the correct classification results, in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Labisch teaches: […] how many data sets have correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] how many data sets have correct classification results […] is disclosed, because Labisch teaches detecting the classifications produced by each model, comparing those classifications with the underlying correct or true classifications, and generating a data set or histogram indicating the number of correct classifications. Labisch’s cases correspond to the data sets, the classifications matching the underlying correct or true classifications correspond to the correct classification results, and the number of correct classifications corresponds to how many data sets have correct classification results) […] having a relatively small number of data sets with the correct classification results (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] having a relatively small number of data sets with the correct classification results is disclosed, because Labisch teaches determining and comparing the number of cases in which classification results match the underlying correct or true classifications. Labisch’s cases correspond to the data sets, the matching classifications correspond to the correct classification results, and a comparatively lower number of correct classifications shown in the generated data set or histogram corresponds to a relatively small number of data sets with correct classification results) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chung with Labisch by using Labisch’s correct classification counting technique as the performance index for each of Chung’s data slices. Chung teaches classifying data sets for each of a plurality of attribute-value patterns different from each other, acquiring a performance value for each attribute-value pattern, and identifying one or more attribute-value patterns having relatively poor classification results. Labisch further teaches determining how many data sets have correct classification results by comparing the classification results obtained by a model with the underlying correct or true classifications and generating statistical information indicating the number of correct and incorrect classifications. Therefore, a POSITA would have been motivated to incorporate Labisch’s correct-classification counting technique into Chung’s process so that, for each of Chung’s plurality of attribute-value patterns different from each other, an index value indicating how many data sets have correct classification results could be acquired. The acquired index values could then be used to identify one or more first attribute-value patterns having a relatively small number of data sets with correct classification results, thereby providing a direct and reliable measure for identifying the attribute-value patterns on which the classifier performs poorly (Labisch, Par. [0031] “This can make it possible, in particular, to determine conditional confidences, which allow a particularly reliable assessment of the correctness of a classification”, & Par. [0032], “when performing the diagnosis method in accordance with the disclosed embodiments of the invention, the data set can be updated accordingly, i.e., the occurrence of incorrect or correct classifications incremented. Processing of an increasing volume of data makes it possible to improve the reliability of the determined confidence, therefore”) Chung combined with Labisch does not explicitly teach in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Chatterjee teaches: in a case of classifying a target data set, determining whether at least any one of […] matches the target data set (Chatterjee, Par. [0060], “After the classifier training is complete, a new observation record is received and a prediction 808 is provided for it. This post-training observation record 806 has a set of attribute values 807”, & Par. [0061], “the post-training observation record 806, the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, thus in a case of classifying a target data set, determining whether at least any one of […] matches the target data set is disclosed, because Chatterjee teaches receiving and classifying a new observation record, examining the identified rules to determine whether any rule is applicable, and determining whether the attribute values of the new observation record match the attribute predicates of at least one rule. Chatterjee’s new observation record corresponds to the target data set, and Chatterjee’s attribute predicate rules correspond to the identified attribute value patterns) and outputting a result of the determining (Chatterjee, Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”, & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a “no explanation is currently available” message and the updating of metrics or statistics regarding the effectiveness of the explainer”, thus and outputting a result of the determining is disclosed, because Chatterjee teaches outputting a matching rule as an explanation when the attribute values of the post-training observation record match the predicates of at least one rule, and outputting a “no explanation is currently available” message when the attribute values do not match any rule predicates. Chatterjee’s provided matching rule or no-explanation message corresponds to the output result of determining whether an identified attribute value pattern matches the target data set) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chung and Labisch with Chatterjee by applying Chatterjee’s attribute predicate matching process to the first attribute value patterns identified using Chung’s performance analysis and Labisch’s correct classification counts. Chung and Labisch collectively teach acquiring, for each of a plurality of attribute value patterns, an index value indicating how many data sets have correct classification results and identifying one or more first attribute value patterns having a relatively small number of data sets with correct classification results. Chatterjee further teaches, after classifying a new observation record, examining attribute predicate rules to determine whether any rule matches the attribute values of the new observation record and outputting either the matching rule or an indication that no rule applies. Therefore, a POSITA would have been motivated to compare a target data set with the identified first attribute value patterns and output the result of the comparison so that the classification result of the target data set could be evaluated based on whether the target data set matches a previously identified attribute value pattern associated with relatively few correct classification results. This would allow the system to inform the user whether the target data set falls within a pattern for which the classifier has demonstrated relatively poor classification performance (Chatterjee, Par. [0061], “The explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”) Regarding Claim 13, Chung teaches: acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier (Chung, Page 2 – Section II, “A slice S is a subset of examples in D with common features and can be described as a conjunction of the common feature-value pairs j Fj op vj where the Fj’s are distinct (e.g., country = DE ∧ gender = Male), and op can be one of =, . For numeric features, we can discretize their values (e.g., quantiles or equiheight bins) and generate ranges so that they are effectively categorical features (e.g., age = [20,30))” & Page 3 – Section II, “We also assume a classification loss function ψ(S,h) that returns a performance score for a set of examples by comparing h’s prediction h(x(i) F ) with the true label y(i). A common classification loss function is logarithmic loss (log loss)”, thus acquiring an index value, the index value indicating […] obtained by classifying data sets for each of a plurality of attribute value patterns different from each other with a classifier is disclosed, because Chung teaches defining respective data slices using different combinations of feature values and obtaining a classification-performance score for the examples in each slice by comparing the classifier’s predictions with the true labels. Chung’s slices defined by conjunctions of feature value pairs correspond to the plurality of attribute value patterns different from each other, Chung’s examples within each slice correspond to the data sets, and Chung’s classification loss or performance score calculated for each slice corresponds to the acquired index value) identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart. The counterpart slice serves as a reference to which we measure how problematic is S, and the definition depends on the problem in hand”, & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h (i.e., most erroneous examples are contained in S and not in S), which should deserve the user’s attention for deeper analysis”, thus identifying, based on the acquired index values, one or more of first attribute value patterns among the plurality of attribute value patterns, each of the one or more of the first attribute value patterns being an attribute value pattern […] is disclosed, because Chung teaches comparing the classification loss value of each slice with that of its counterpart and identifying a problematic slice having a higher concentration of classification errors. Chung’s classification loss values correspond to the acquired index values, Chung’s slices defined by common feature value pairs correspond to the plurality of attribute value patterns, and Chung’s identified problematic slices correspond to the one or more first attribute value patterns) […] the identified one or more of first attribute value patterns […] (Chung, Page 3 – Section II, “We define a slice to be problematic if the classification loss function takes vastly different values between the slice and its counterpart,” & Page 3 – Section II, “This effectively allows us to identify S with a higher error concentration for h,” thus […] the identified one or more of first attribute value patterns […] is disclosed, because Chung teaches identifying one or more problematic slices having a higher concentration of classification errors. Chung’s identified problematic slices correspond to the identified one or more first attribute value patterns, and the common feature value pairs defining each problematic slice correspond to the attribute values forming each first attribute value pattern) Chung does not explicitly teach a memory, a processor, […] how many data sets have correct classification results […], […] a relatively small number of data sets with the correct classification results, in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Labisch teaches: a memory (Labisch, Claim 12, “memory;”, thus a memory is disclosed) a processor (Labisch, Claim 12, “ a processor”, thus a processor is disclosed) […] how many data sets have correct classification results […] (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] how many data sets have correct classification results […] is disclosed, because Labisch teaches detecting the classifications produced by each model, comparing those classifications with the underlying correct or true classifications, and generating a data set or histogram indicating the number of correct classifications. Labisch’s cases correspond to the data sets, the classifications matching the underlying correct or true classifications correspond to the correct classification results, and the number of correct classifications corresponds to how many data sets have correct classification results) […] having a relatively small number of data sets with the correct classification results (Labisch, Par. [0030], “for each of the models, the number of classifications resulting from the application of the appropriate model can be detected and compared with the number of cases in which the classification by the model matches the underlying correct or true classification, in particular are related to each other. In other words, a data set, for instance, in the form of a histogram, can be generated from which for each of the models, the number of correct classifications and/or the number of incorrect classifications emerges”, thus […] having a relatively small number of data sets with the correct classification results is disclosed, because Labisch teaches determining and comparing the number of cases in which classification results match the underlying correct or true classifications. Labisch’s cases correspond to the data sets, the matching classifications correspond to the correct classification results, and a comparatively lower number of correct classifications shown in the generated data set or histogram corresponds to a relatively small number of data sets with correct classification results) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chung with Labisch by using Labisch’s correct classification counting technique as the performance index for each of Chung’s data slices. Chung teaches classifying data sets for each of a plurality of attribute-value patterns different from each other, acquiring a performance value for each attribute-value pattern, and identifying one or more attribute-value patterns having relatively poor classification results. Labisch further teaches determining how many data sets have correct classification results by comparing the classification results obtained by a model with the underlying correct or true classifications and generating statistical information indicating the number of correct and incorrect classifications. Therefore, a POSITA would have been motivated to incorporate Labisch’s correct-classification counting technique into Chung’s process so that, for each of Chung’s plurality of attribute-value patterns different from each other, an index value indicating how many data sets have correct classification results could be acquired. The acquired index values could then be used to identify one or more first attribute-value patterns having a relatively small number of data sets with correct classification results, thereby providing a direct and reliable measure for identifying the attribute-value patterns on which the classifier performs poorly (Labisch, Par. [0031] “This can make it possible, in particular, to determine conditional confidences, which allow a particularly reliable assessment of the correctness of a classification”, & Par. [0032], “when performing the diagnosis method in accordance with the disclosed embodiments of the invention, the data set can be updated accordingly, i.e., the occurrence of incorrect or correct classifications incremented. Processing of an increasing volume of data makes it possible to improve the reliability of the determined confidence, therefore”) Chung combined with Labisch does not explicitly teach in a case of classifying a target data set, determining whether at least any one of […] matches the target data set, and outputting a result of the determining. However, Chatterjee teaches: in a case of classifying a target data set, determining whether at least any one of […] matches the target data set (Chatterjee, Par. [0060], “After the classifier training is complete, a new observation record is received and a prediction 808 is provided for it. This post-training observation record 806 has a set of attribute values 807”, & Par. [0061], “the post-training observation record 806, the explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, thus in a case of classifying a target data set, determining whether at least any one of […] matches the target data set is disclosed, because Chatterjee teaches receiving and classifying a new observation record, examining the identified rules to determine whether any rule is applicable, and determining whether the attribute values of the new observation record match the attribute predicates of at least one rule. Chatterjee’s new observation record corresponds to the target data set, and Chatterjee’s attribute predicate rules correspond to the identified attribute value patterns) and outputting a result of the determining (Chatterjee, Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”, & Par. [0063], “if the attributes values 807 of the post-training observation record do not match any of the predicates for which rules have been generated, this may also result in a “no explanation is currently available” message and the updating of metrics or statistics regarding the effectiveness of the explainer”, thus and outputting a result of the determining is disclosed, because Chatterjee teaches outputting a matching rule as an explanation when the attribute values of the post-training observation record match the predicates of at least one rule, and outputting a “no explanation is currently available” message when the attribute values do not match any rule predicates. Chatterjee’s provided matching rule or no-explanation message corresponds to the output result of determining whether an identified attribute value pattern matches the target data set) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chung and Labisch with Chatterjee by applying Chatterjee’s attribute predicate matching process to the first attribute value patterns identified using Chung’s performance analysis and Labisch’s correct classification counts. Chung and Labisch collectively teach acquiring, for each of a plurality of attribute value patterns, an index value indicating how many data sets have correct classification results and identifying one or more first attribute value patterns having a relatively small number of data sets with correct classification results. Chatterjee further teaches, after classifying a new observation record, examining attribute predicate rules to determine whether any rule matches the attribute values of the new observation record and outputting either the matching rule or an indication that no rule applies. Therefore, a POSITA would have been motivated to compare a target data set with the identified first attribute value patterns and output the result of the comparison so that the classification result of the target data set could be evaluated based on whether the target data set matches a previously identified attribute value pattern associated with relatively few correct classification results. This would allow the system to inform the user whether the target data set falls within a pattern for which the classifier has demonstrated relatively poor classification performance (Chatterjee, Par. [0061], “The explainer may examine the rule set 804, e.g., in order of decreasing rank, to determine whether any of the rules in its rule set is applicable or not”, & Par. [0062], “If the attribute values 807 match (or overlap) with the attribute predicates 802 of at least one rule, and the prediction 808 matches the implication in that rule, such a rule may be provided to the client as an explanation for the prediction”) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. WO 2018079020 A1 is pertinent because it teaches an information processing system for evaluating or predicting machine learning performance based on the state of labeled learning data. The reference therefore concerns analyzing data associated with a machine learning process and generating information indicative of expected classification or learning performance. Because applicant’s disclosure similarly concerns evaluating classification performance based on data sets and classification results, including identifying data patterns associated with relatively favorable or unfavorable classification performance, the reference is relevant to the claimed invention but is not relied upon in the rejection. The reference does not clearly teach acquiring correct classification counts for each of a plurality of attribute value patterns, matching a target data set with identified first and second attribute value patterns, or selecting a classifier based on such matching. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Mar 08, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month