Prosecution Insights
Last updated: August 17, 2026
Application No. 18/065,296

DOMAIN KNOWLEDGE-BASED EVALUATION OF MACHINE LEARNING MODELS

Final Rejection §101§102§103
Filed
Dec 13, 2022
Examiner
ROHD, BENJAMIN MATTHEW
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
2 (Final)
0%
Grant Probability
At Risk
3-4
OA Rounds
7m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 2 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
21 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
24.9%
-15.1% vs TC avg
§103
49.2%
+9.2% vs TC avg
§102
9.6%
-30.4% vs TC avg
§112
15.8%
-24.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 2 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION This office action is in response to amendments filed on 01/26/2026. Claims 1, 3, 7, 9-11, 16, and 18-20 have been amended. Claims 1-20 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Claim Objections: In light of applicant’s amendments to the claims (pg. 2-6), the objections to the claims have been withdrawn. Rejections Under 35 USC § 101: Applicant's arguments regarding the rejections under 35 USC § 101 have been fully considered but they are not persuasive. Applicant argues (pg. 7-8) that the claims are patent eligible because they are directed to a technical improvement in the field of machine learning. Specifically, applicant asserts that the incorporation of domain knowledge via comparison between expert input and model explanations enables evaluation of the scientific validity of ML models and thereby improves model selection beyond consideration of accuracy alone. In response, examiner respectfully notes that, per MPEP 2106.05(a), “It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements… In addition, the improvement can be provided by the additional element(s) in combination with the recited judicial exception.” In this case, as acknowledged by applicant, the alleged improvement of “domain knowledge-based evaluation of machine learning (ML) models for a target subjection” is achieved by the “comparing” step, which, as can be seen in the rejection below, is considered a mental process. Therefore, the improvement is provided by the judicial exception, and is thus not sufficient to amount to significantly more than the judicial exception. The rejections under 35 USC § 101 have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Prior Art Rejections: Applicant's arguments regarding the prior art rejections have been fully considered but they are not persuasive. Applicant argues (pg. 8-9) that while DeYoung discloses prediction rationales provided by the ML models, DeYoung does not disclose the claimed “model explanations” which “include a set of identified important features.” Examiner respectfully disagrees. DeYoung teaches that the rationales include token (feature) importance: “We would also like to evaluate the faithfulness of continuous importance scores assigned to tokens by models. Here we adopt a simple approach for this. We convert soft scores over features s i provided by a model into discrete rationales r i …” (DeYoung, pg. 6, section 4.2). Applicant argues (pg. 8-9) that DeYoung does not disclose the claimed “target subject” because DeYoung does not suggest that the corpora of NLP datasets represent a target subject or a single domain. Examiner respectfully notes that “target subject” is not a known term of art, nor is it further defined in the specification. Under its broadest reasonable interpretation, the “target subject” simply refers to the data on which the machine learning models operate. Therefore, the corpora of NLP data disclosed by DeYoung falls within the broadest reasonable interpretation of the claimed “target subject.” Applicant argues (pg. 9) that while DeYoung discloses measuring agreement between determined and human-provided rationales, DeYoung does not disclose the claimed “comparing… by determining whether at least one of the known important features matches at least one identified important feature of the model explanations.” Examiner respectfully disagrees. DeYoung teaches measuring similarity between “predicted and reference rationales” by determining “the size of the overlap of the tokens they cover” (i.e. whether identified important features match known important features) (DeYoung, pg. 5, section 4.1). The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Claim Interpretation Claim 20 recites “A computer program product for domain knowledge-based evaluation of machine learning models for a target subject, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor”. According to the specification paragraph [0020] describes “A computer program product” as “A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any set of one, or more, storage media (also called “mediums”) collectively included in a set of one, or more, storage devices that collectively include machine readable code corresponding to instructions and/or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can retain and store instructions for use by a computer processor. Without limitation, the computer readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these mediums include: diskette, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device (such as punch cards or pits/lands formed in a major surface of a disc) or any suitable combination of the foregoing. A computer readable storage medium, as that term is used in the present disclosure, is not to be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through a fiber optic cable, electrical signals communicated through a wire, and/or other transmission media.” Therefore, “A computer program product” in claim 20 will be interpreted as a non-transitory medium. The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: accessing, by a model explanation module… in claim 1; receiving, by a domain expert input module… in claim 1; comparing, by a comparing module… in claim 1; scoring, by a scoring module… in claim 1. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Per specification paragraph 0090, the claimed “modules” are interpreted as referring to software implementations of the method steps with which they are associated: “The ML model analysis module 400 includes modules in the form of software units providing the functionality of the described method.” If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because, in light of the interpretation of the claimed “modules” under 35 U.S.C. 112(f) based on the corresponding structure in specification paragraph 0090, the claimed invention is directed to software per se – see MPEP 2106.03(I). Assuming the software per se rejection above were remedied, Claims 1-20 are additionally rejected under 35 U.S.C. 101 because the claims are directed to an abstract idea without significantly more. Regarding claim 1: Subject Matter Eligibility Analysis Step 1: Claim 1 recites “A computer-implemented method for domain knowledge-based evaluation of machine learning (ML) models for a target subject” which is a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 1 recites the steps: “upon accessing the model explanations and receiving the domain expert input, automatically comparing, by a comparing module, the domain expert input with the model explanations to evaluate consensus by determining whether at least one of the known important features matches at least one identified important feature of the model explanations”: Comparing model explanations and domain expert input to evaluate consensus by determining matching important features can practically be performed in the human mind or with the aid of pen and paper. Hence, this is a mental process. “scoring, by a scoring module, the ML models based on the evaluated consensus”: Scoring ML models based on the consensus determined in the previous step can practically be performed in the human mind or with the aid of pen and paper. Hence, this is also a mental process. Claim 1 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 1 recites the additional elements: “accessing, by a model explanation module, model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). “receiving, by a domain expert input module, a domain expert input including a set of known important features for the prediction task in the target subject”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because it recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). Thus, claim 1 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional elements in claim 1 do not provide significantly more than the abstract ideas themselves, taken alone and in combination because: “accessing, by a model explanation module, model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). “receiving, by a domain expert input module, a domain expert input including a set of known important features for the prediction task in the target subject”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 1 is subject-matter ineligible. Regarding claim 2: Subject Matter Eligibility Analysis Step 1: Claim 2 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 2 recites the same mental concepts as claim 1, therefore claim 2 recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: In addition to the additional elements in claim 1, claim 2 recites: “wherein the model explanation includes the set of identified important features and an associated effect of each feature on the target subject, and the domain expert input includes the set of known important features and an associated effect of each feature on the target subject, wherein the associated effect of a feature on the target subject includes a directional agreement with the target subject”: This element does not integrate the abstract ideas into a practical application because the element further expands on the limitations “accessing a model explanation of each of a plurality of machine learning models for the target subject, wherein a model explanation includes a set of identified important features” and “receiving a domain expert input including a set of known important features for the target subject” in claim 1, hence this element is an insignificant extra solution activity of data transmission (MPEP 2106.05(g). Claim 2 thus is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional element in claim 2 does not provide significantly more than the abstract ideas themselves taken alone and in combination because: “wherein the model explanation includes the set of identified important features and an associated effect of each feature on the target subject, and the domain expert input includes the set of known important features and an associated effect of each feature on the target subject, wherein the associated effect of a feature on the target subject includes a directional agreement with the target subject”: This element does not integrate the abstract ideas into a practical application because the element further expands on the limitations “accessing a model explanation of each of a plurality of machine learning models for the target subject, wherein a model explanation includes a set of identified important features” and “receiving a domain expert input including a set of known important features for the target subject” in claim 1, hence the element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 2 is subject-matter ineligible. Regarding claim 3: Subject Matter Eligibility Analysis Step 1: Claim 3 is a process as in claim 2. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 2, claim 3 recites: “determining a match of the directional agreement of the features with the target subject”: Determining a match in directional agreement of features and the target subject can practically be performed in the human mind or with the aid of pen and paper. Hence, this is a mental process. Therefore, claim 3 recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 3 recites the same additional elements as claim 2, thus claim 3 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 3 recites the same additional elements as claim 2 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 3 is subject-matter ineligible. Regarding claim 4: Subject Matter Eligibility Analysis Step 1: Claim 4 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 4 recites the same mental concepts as claim 1, therefore claim 4 recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: In addition to the additional elements in claim 1, claim 4 recites: “receiving a direct input from a domain expert via a user interface”: This element does not integrate the abstract ideas into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g). Claim 4 thus is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional element in claim 4 does not provide significantly more than the abstract ideas themselves taken alone and in combination because: “receiving a direct input from a domain expert via a user interface”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 4 is subject-matter ineligible. Regarding claim 5: Subject Matter Eligibility Analysis Step 1: Claim 5 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 1, claim 5 mentions: “evaluating one or more linked domain expert resources”: Evaluating domain expert resources can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. “extracting important features from the domain expert resources for the target subject”: Extracting features from domain expert resources for the target subject can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 5 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 5 recites the same additional elements as claim 1, thus claim 5 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 5 recites the same additional elements as claim 1 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 5 is subject-matter ineligible. Regarding claim 6: Subject Matter Eligibility Analysis Step 1: Claim 6 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 1, claim 6 mentions: “resolving different specific features within a same genus”: Resolving features within a genus can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 6 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 6 recites the same additional elements as claim 1, thus claim 6 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 6 recites the same additional elements as claim 1 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 6 is subject-matter ineligible. Regarding claim 7: Subject Matter Eligibility Analysis Step 1: Claim 7 is a process as in claim 6. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 6, claim 7 mentions: “analyzing edges in a network to determine significant association between nodes representing specific features”: Analyzing edges in a network to determine associations between nodes can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 7 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: In addition to the additional elements in claim 6, claim 7 recites: “using normalized abundance counts from the ML model training as input for network construction”: This element does not integrate the abstract ideas into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g). Claim 7 thus is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional element in claim 7 does not provide significantly more than the abstract ideas themselves taken alone and in combination because: “using normalized abundance counts from the machine learning model training as input for network construction”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 7 is subject-matter ineligible. Regarding claim 8: Subject Matter Eligibility Analysis Step 1: Claim 8 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 1, claim 8 mentions: “resolving one or more specific features to compare to generic features by accessing additional resources to analyze specific features.”: Resolving features by using additional resources to analyze those features can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 8 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 8 recites the same additional elements as claim 1, thus claim 8 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 8 recites the same additional elements as claim 1 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 8 is subject-matter ineligible. Regarding claim 9: Subject Matter Eligibility Analysis Step 1: Claim 9 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 1, claim 9 mentions: “scoring the plurality of ML models by applying weightings for aspects of the comparison of the domain expert input with the plurality of model explanations.”: Applying weightings during the comparison in claim 1 in order to score the ML models can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 9 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 9 recites the same additional elements as claim 1, thus claim 9 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 9 recites the same additional elements as claim 1 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 9 is subject-matter ineligible. Regarding claim 10: Subject Matter Eligibility Analysis Step 1: Claim 10 is a process as in claim 1. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 1, claim 10 mentions: “selecting a best ML model according to a selection strategy from a ranking of a plurality of ML models based on the scoring.”: Selecting the best ML model from a score-based ranking of ML models can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. Claim 10 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 10 recites the same additional elements as claim 1, thus claim 10 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 10 recites the same additional elements as claim 1 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 10 is subject-matter ineligible. Regarding claim 11: Subject Matter Eligibility Analysis Step 1: Claim 11 recites “A system for analyzing machine learning models for domain knowledge-based evaluation of machine learning (ML) models for a target subject, the system comprising: a processor and a memory configured to provide computer program instructions to the processor to execute the function of code modules….” which is an article of manufacture, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 11 recites the steps: “analyzing machine learning models for domain knowledge-based evaluation of machine learning (ML) models for a target subject”: Analyzing ML models for a target subject can practically be performed in the human mind or with the aid of pen and paper. Hence this is a mental process. “upon accessing the model explanations and receiving the domain expert input, automatically comparing, by a comparing module, the domain expert input with the model explanations to evaluate consensus by determining whether at least one of the known important features matches at least one identified important feature of the model explanations”: Comparing model explanations and domain expert input to evaluate consensus by determining matching important features can practically be performed in the human mind or with the aid of pen and paper. Hence, this is a mental process. “scoring, by a scoring module, the ML models based on the evaluated consensus”: Scoring ML models based on the consensus determined in the previous step can practically be performed in the human mind or with the aid of pen and paper. Hence, this is also a mental process. Claim 11 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 11 recites the additional elements: “a processor and a memory configured to provide computer program instructions to the processor to execute the function of code modules”: This element does not integrate the abstract ideas into a practical application because it recites generic computing components (MPEP 2106.05(f)). “accessing, by a model explanation module, model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). “receiving, by a domain expert input module, a domain expert input including a set of known important features for the prediction task in the target subject”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because it recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). Thus, claim 11 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional elements in claim 11 do not provide significantly more than the abstract ideas themselves, taken alone and in combination because: “a processor and a memory configured to provide computer program instructions to the processor to execute the function of code modules”: This element recites generic computing components (MPEP 2106.05(f)). “accessing, by a model explanation module, model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). “receiving, by a domain expert input module, a domain expert input including a set of known important features for the prediction task in the target subject”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 11 is subject-matter ineligible. Regarding claim 12: Claim 12 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 2. Regarding claim 13: Claim 13 is an article of manufacture as in claim 11 and is rejected for the same rationale as claim 3. Regarding claim 14: Claim 14 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 4. Regarding claim 15: Claim 15 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 5. Regarding claim 16: Claim 16 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 7. Regarding claim 17: Claim 17 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 8. Regarding claim 18: Claim 18 is an article of manufacture as in claim 11 and is rejected for the same reason as claim 9. Regarding claim 19: Subject Matter Eligibility Analysis Step 1: Claim 19 is a process as in claim 11. Subject Matter Eligibility Analysis Step 2A Prong 1: In addition to the mental concepts in claim 11, claim 19 mentions: “configuring analysis metrics for the ML models.”: Configuring analysis metrics for ML models can practically be performed in the human mind or with the aid of pen and paper. Hence, this is a mental process. Claim 19 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 19 recites the same additional elements as claim 11, thus claim 19 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: Since claim 19 recites the same additional elements as claim 11 and since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 19 is subject-matter ineligible. Regarding claim 20: Subject Matter Eligibility Analysis Step 1: Claim 20 recites “A computer program product for domain knowledge-based evaluation of machine learning (ML) models for a target subject, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor” and as stated above, since the “computer program product” is directed to a non-transitory medium, it is directed to a machine, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 20 recites the steps: “upon accessing the model explanations and receiving the domain expert input, automatically compare the domain expert input with the model explanations to evaluate consensus by determining whether at least one of the known important features matches at least one identified important feature of the model explanations”: Comparing model explanations and domain expert input to evaluate consensus by determining matching important features can practically be performed in the human mind or with the aid of pen and paper. Hence, this is a mental process. “score the ML models based on the evaluated consensus”: Scoring ML models based on the consensus determined in the previous step can practically be performed in the human mind or with the aid of pen and paper. Hence, this is also a mental process. Claim 20 therefore recites abstract ideas. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 20 recites the additional elements: “the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because the element recites generic computing components (MPEP 2106.05(f)). “access model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because the element recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). “receive a domain expert input including a set of known important features for the prediction task in the target subject”: This element does not integrate the abstract ideas from Step 2A Prong 1 into a practical application because it recites an insignificant extra solution activity of data transmission (MPEP 2106.05(g)). Thus, claim 20 is directed to the abstract ideas. Subject Matter Eligibility Analysis Step 2B: The additional elements in claim 20 do not provide significantly more than the abstract ideas themselves, taken alone and in combination because: “the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor”: This element recites generic computing components (MPEP 2106.05(f)). “access model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). “receive a domain expert input including a set of known important features for the prediction task in the target subject”: This element mentions a well-understood, routine, and conventional activity of “receiving or transmitting data over a network” (MPEP 2106.05(d)(I), Intellectual Ventures v. Symantec, 838 F.3d 1307, 1321; 120 USPQ2d 1353, 1362 (Fed. Cir. 2016) [utilizing an intermediary computer to forward information]). Since there is no nexus between additional elements that could cause the combination to provide an inventive concept, claim 20 is subject-matter ineligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-5, 9, and 10 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by DeYoung et al.’s “ERASER: A Benchmark to Evaluate Rationalized NLP Models”. Regarding claim 1, DeYoung et al. teaches “A computer-implemented method for domain knowledge-based evaluation of machine learning (ML) models for a target subject” (DeYoung et al. Section 1 “In sum, we introduce the ERASER benchmark (www.eraserbenchmark.com), a unified set of diverse NLP datasets (these are repurposed and augmented from existing corpora,1 including sentiment analysis, Natural Language Inference, and QA tasks, among others) in a standardized format featuring human rationales for decisions, along with starter code and tools, baseline models, and standardized (initial) metrics for rationales.”; Appendix Section C “This model was implemented using the AllenNLP library”; “benchmark” and “AllenNLP library” corresponds to “A computer-implemented method”; “standardized (initial) metrics for rationales” corresponds “domain knowledge-based evaluation”; “baseline models” corresponds to “machine learning models”; “diverse NLP datasets” corresponds to “a target subject”) “the computer-implemented method comprising:” “accessing, by a model explanation module, model explanations respectively corresponding to individual ML models of a plurality of ML models trained in the target subject, wherein the model explanations include a set of identified important features for a prediction task performed by the individual ML models” (DeYoung et al., Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs. All datasets in ERASER include such rationales, explicitly marked by human annotators. By definition, rationales should be sufficient to make predictions, but they may not be comprehensive. Therefore, for some datasets, we have also collected comprehensive rationales (in which all evidence supporting an output has been marked) on test instances. The ‘quality’ of extracted rationales will depend on their intended use. Therefore, we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER. We find that no single ‘off-the-shelf’ architecture is readily adaptable to datasets with very different instance lengths and associated rationale snippets.” DeYoung et al., Section 4.2 “We would also like to evaluate the faithfulness of continuous importance scores assigned to tokens by models. Here we adopt a simple approach for this. We convert soft scores over features s i provided by a model into discrete rationales r i …”; a “rationale” “provided by a model” corresponds to “a model explanation”; “baseline models” corresponds to “a plurality of machine learning models”; “the corpora” corresponds to “target subject”; since the rationales are “snippets that supports outputs” and “are sufficient to make predictions”, they correspond to “a set of identified important features for a prediction task performed by the individual ML models”; implementation of the method in a computing environment, as described in regard to the claim’s preamble, corresponds to “a model explanation module” (see the interpretation of “module” under 35 USC 112(f) above).); “receiving, by a domain expert input module, a domain expert input including a set of known important features for the prediction task in the target subject” (DeYoung et al., Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs. All datasets in ERASER include such rationales, explicitly marked by human annotators. By definition, rationales should be sufficient to make predictions, but they may not be comprehensive. Therefore, for some datasets, we have also collected comprehensive rationales (in which all evidence supporting an output has been marked) on test instances. The ‘quality’ of extracted rationales will depend on their intended use. Therefore, we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER. We find that no single ‘off-the-shelf’ architecture is readily adaptable to datasets with very different instance lengths and associated rationale snippets”; “human-provided rationales” correspond to “receiving a domain expert input including a set of known important features for the prediction task”; “the corpora” corresponds to “the target subject”; implementation of the method in a computing environment, as described in regard to the claim’s preamble, corresponds to “a domain expert input module” (see the interpretation of “module” under 35 USC 112(f) above).); “upon accessing the model explanation and receiving the domain expert input, automatically comparing, by a comparing module, the domain expert input with the model explanations to evaluate consensus by determining whether at least one of the known important features matches at least one identified important feature of the model explanations” (DeYoung et al., Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs. All datasets in ERASER include such rationales, explicitly marked by human annotators. By definition, rationales should be sufficient to make predictions, but they may not be comprehensive. Therefore, for some datasets, we have also collected comprehensive rationales (in which all evidence supporting an output has been marked) on test instances. The ‘quality’ of extracted rationales will depend on their intended use. Therefore, we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER. We find that no single ‘off-the-shelf’ architecture is readily adaptable to datasets with very different instance lengths and associated rationale snippets.” Section 4.1: “For the discrete case, measuring exact matches between predicted and reference rationales is likely too harsh. We thus consider more relaxed measures. These include Intersection-Over-Union (IOU), borrowed from computer vision (Everingham et al., 2010), which permits credit assignment for partial matches. We define IOU on a token level: for two spans, it is the size of the overlap of the tokens they cover divided by the size of their union”; “initial set of metrics to evaluate rationales…includes measures of agreement with human-provided rationales” corresponds to “comparing… the domain expert input with the model explanations to evaluate consensus”; measuring similarity between “predicted and reference rationales” by determining “the size of the overlap of the tokens they cover” corresponds to “determining whether at least one of the known important features matches at least one identified important feature of the model explanations”; implementation of the method in a computing environment, as described in regard to the claim’s preamble, corresponds to “a comparing module” (see the interpretation of “module” under 35 USC 112(f) above).); and “scoring, by a scoring module, the ML models based on the evaluated consensus” (DeYoung et al., Section 1 “Therefore, we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness.”; Table 3 “Performance of models that perform hard rationale selection. All models are supervised at the rationale level except for those marked with (u), which learn only from instance-level supervision; † denotes cases in which rationale training degenerated due to the REINFORCE style training. Perf. is accuracy (CoS-E) or macro-averaged F1 (others). Bert-To-Bert for CoS-E and e-SNLI uses a token classification objective. BertTo-Bert CoS-E uses the highest scoring answer.”; Table 3 ranks the “performance of models”, thus it corresponds to “scoring the machine learning models”; “evaluate rationales…includes measures of agreement with human-provided rationales” corresponds to “based on the evaluated consensus”; implementation of the method in a computing environment, as described in regard to the claim’s preamble, corresponds to “a scoring module” (see the interpretation of “module” under 35 USC 112(f) above).) Regarding claim 2, the rejections of claim 1 are incorporated. DeYoung et al. further mentions: “the model explanation includes the set of identified important features and an associated effect of each feature on the target subject, and the domain expert input includes the set of known important features and an associated effect of each feature on the target subject, wherein the associated effect of a feature on the target subject includes a directional agreement with the target subject” (DeYoung et al., Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs. All datasets in ERASER include such rationales, explicitly marked by human annotators. By definition, rationales should be sufficient to make predictions, but they may not be comprehensive. Therefore, for some datasets, we have also collected comprehensive rationales (in which all evidence supporting an output has been marked) on test instances. The ‘quality’ of extracted rationales will depend on their intended use. Therefore, we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER. We find that no single ‘off-the-shelf’ architecture is readily adaptable to datasets with very different instance lengths and associated rationale snippets”; since the “rationales” are both “snippets that support outputs” and are “sufficient to make predictions” for “the corpora”, the rationales “includes the set of identified important features and an associated effect of each feature on the target subject”; “measures of agreement” corresponds to “directional agreement”) Regarding claim 3, the rejections claim 2 are incorporated. DeYoung et al. further recites: “determining a match of the directional agreement of the features with the target subject” (DeYoung et al., Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs…we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER.”; “rationales provided by a model” corresponds to “identified important features”; “human-provided rationales” correspond to “known important features”; “the corpora” corresponds to “target subject”; “evaluate rationales…includes measure of agreement with human-provided rationales” correspond to “determining a match”); Regarding claim 4, the rejections of claim 1 are incorporated. DeYoung et al. further mentions: “receiving a direct input from a domain expert via a user interface” (DeYoung et al. Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs. All datasets in ERASER include such rationales, explicitly marked by human annotators.…we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER.”; Appendix “We used the Upwork Platform12 to hire two fluent english speakers to annotate each of the 200 documents in our test set. Workers were paid at rate of USD 8.5 per hour and on average, it took them 5 min to annotate a document. Each annotator was asked to annotate a set of 6 documents and compared against in-house annotations (by authors).” “human annotators” and “two fluent english speakers” correspond to “a domain expert”; “Upwork Platform12” corresponds to “a user interface”) Regarding claim 5, the rejections of claim 1 are incorporated. DeYoung et al. teaches: “evaluating one or more linked domain expert resources” (DeYoung et al. Appendix “Movies. We used the Upwork Platform12 to hire two fluent english speakers to annotate each of the 200 documents in our test set. Workers were paid at rate of USD 8.5 per hour and on average, it took them 5 min to annotate a document. Each annotator was asked to annotate a set of 6 documents and compared against in-house annotations (by authors).”; “compared against in-house annotations” corresponds to “evaluating one or more”; “200 documents in our test set” corresponds to “linked domain expert resources”); and “extracting important features from the domain expert resources for the target subject” (DeYoung et al. Section 1 “In ERASER we focus specifically on rationales, i.e., snippets that support outputs…we propose an initial set of metrics to evaluate rationales that are meant to measure different varieties of ‘interpretability’. Broadly, this includes measures of agreement with human-provided rationales, and assessments of faithfulness. The latter aim to capture the extent to which rationales provided by a model in fact informed its predictions… We implement baseline models and report their performance across the corpora in ERASER.”; “rationales provided by a model” correspond to “extracting important features”; “the corpora in ERASER” corresponds to “the target subject”) Regarding claim 9, the rejections of claim 1 are incorporated. DeYoung et al. further teaches: “scoring the plurality of ML models by applying weightings for aspects of the comparison of the domain expert inputs with the plurality of model explanations” (DeYoung et al. Section 5 “Our focus in this work is primarily on the ERASER benchmark itself, rather than on any particular model(s). But to establish a starting point for future work, we evaluate several baseline models across the corpora in ERASER.8 We broadly classify these into models that assign ‘soft’ (continuous) scores to tokens, and those that perform a ‘hard’ (discrete) selection over inputs.”; Section 6 “In Table 3 we evaluate models that perform discrete selection of rationales. We view these as inherently faithful, because by construction we know which snippets the decoder used to make a prediction. Therefore, for these methods we report only metrics that measure agreement with human annotations…In Table 4 we report metrics for models that assign continuous importance scores to individual tokens…Both simple gradient and LIME-based scoring yield more comprehensive rationales than attention weights, consistent with prior work. Attention fares better in terms of AUPRC — suggesting better agreement with human rationales….”; “report metrics for models” correspond to “scoring the plurality of machine learning models”; since “simple gradient”, “LIME-based scoring” and “attention weights” scorings were utilized to benchmark the models, this corresponds to “applying weightings”; since the scorings on the rationales by the models were performed to “measure agreement with human annotations”, this corresponds to “comparison of the domain expert inputs with the plurality of model explanations”); Regarding claim 10, the rejections of claim 1 are incorporated. DeYoung et al. further mentions: “selecting the best ML model according to a selection strategy from a ranking of a plurality of ML models based on the scoring” (DeYoung et al. Section 6 “In Table 3 we evaluate models that perform discrete selection of rationales. We view these as inherently faithful, because by construction we know which snippets the decoder used to make a prediction.10 Therefore, for these methods we report only metrics that measure agreement with human annotations… We observe that for the “rationalizing” model of Lei et al. (2016), exploiting rationale-level supervision often (though not always) improves agreement with human-provided rationales”; “‘rationalizing’ model of Lei et al.” corresponds to “selecting the best machine learning model”; “metrics that measure agreement with human annotation” corresponds to “a selection strategy”; since Table 3 displays performance of baseline models, it corresponds to “a ranking”) Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 6-8 and 11-20 are rejected under 35 U.S.C. 103 as being unpatentable over DeYoung et al.’s “ERASER: A Benchmark to Evaluate Rationalized NLP Models” in view of Bhusan et al. (US20230067976A1). Regarding claim 6, the rejections of claim 1 are incorporated. DeYoung et al. fails to teach: “resolving different specific features within a same genus” However, Bhusan et al. teaches: “resolving different specific features within a same genus” (Bhusan et al. [0083] “At step 306 of the method 300, bacterial abundance data is obtained from the sample corresponding to the disease using an experimental technique, wherein the bacterial abundance data is used to construct a bacterial taxonomic abundance matrix consisting of abundance information of individual bacterial taxon across the group of patients.”; [0006] “….the bacterial abundance data is used to construct a bacterial taxonomic abundance matrix consisting of abundance information of individual bacterial taxon across the group of patients”; FIG. 3A displays that the bacterial taxonomic abundance matrix is utilized to “construct a first bacterial association network” in Step 308 which is further used to construct a “feature count matrix” in Step 318 for applying a first machine learning model which is labeled as “a first classifier” and FIG. 3B also uses that same “feature count matrix” for generating a “refined association network” in Step 320 then applying a second machine learning model which is labeled as “a second classifier” in Step 322; thus, “construct a bacterial taxonomic abundance matrix consisting of abundance information of individual bacterial taxon” and “construct a first bacterial association network” corresponds to “resolving different specific features”; “bacteria” corresponds to “same genus”) [note: paragraph [0096] defines “resolving different specific features” as “….resolving different specific features within a same genus including: using normalized abundance counts from the machine learning model training as input for network construction”; hence “resolving different specific features” is interpreted to mean constructing a network using abundance information] ; DeYoung et al. and Bhusan et al. are both analogous to the claimed invention because they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the ERASER corpora with specific features within literature as taught by DeYoung et al. with a bacterial abundance matrix with specific features of bacteria as taught by Bhusan et al. Hence, this would be a simple substitution of one known element (ERASER corpora) for another (bacterial abundance data) to obtain predictable results (extracting specific features from the same category). (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results). Regarding claim 7, the rejections of claim 6 are incorporated. Bhusan et al. further mentions: “using normalized abundance counts from the ML model training as input for network construction”(Busan et al. FIG. 3A displays that the normalized bacterial taxonomic abundance matrix is utilized to “construct a first bacterial association network” in Step 308 which is further used to construct a “feature count matrix” in Step 318 for applying a first machine learning model which is labeled as “a first classifier” and FIG. 3B also uses that same “feature count matrix” for generating a “refined association network” in Step 320 then applying a second machine learning model which is labeled as “a second classifier” in Step 322; [0083] describes normalizing the taxonomic abundance matrix which was used in step 308 as “At step 306 of the method 300, bacterial abundance data is obtained from the sample corresponding to the disease using an experimental technique, wherein the bacterial abundance data is used to construct a bacterial taxonomic abundance matrix consisting of abundance information of individual bacterial taxon across the group of patients”; [0084] further describes the data normalization process “A data normalization step is desirable in the next step to remove various sampling and experimental biases. The most common technique of normalization being a total sum scaling or a percentage normalization. However, advanced techniques including rarefaction based or centered log-ratio based transformation may also be used to normalize the abundance matrices. The bacterial taxonomic abundance matrix can then be used to identify relationship patterns between the bacteria (in the matrix) using proxy measures like significant correlations calculated from the matrix…Once the ‘all versus all’ pairwise bacterial associations (or in other words all possible unique pairwise combinations between the available bacteria) are calculated from the bacterial taxonomic abundance matrix, only the significant associations above or below a certain threshold value (which can be the correlation or the probability value of the association) are identified and an edge is connected between that pair of bacteria having a significant association (for example correlations having a probability value<0.05). Upon completion of the task for all the bacteria in the matrix, a bacterial association network is generated.”; [0090] then describes how the abundance matrix and the bacterial association network in step 308 are utilized for machine learning prediction by stating “In the next step 326 of the method 300, sentences in the first list of biomedical text are identified (by a method like sentence tokenization) with probable bacterial associations. At step 328 a table of predicted sentences is created using the first classifier (or sentence classifier trained on the sentence corpus)”; “normalize the abundance matrices” and “The bacterial taxonomic abundance matrix can then be used to identify relationship patterns” correspond to “using normalized abundance counts from the machine learning training”; since the abundance matrix is used to generate “a bacterial association network” corresponds to “input for network construction”; FIG. 3A displays that the bacterial abundance matrix is used to generate “a biomedical text corpus” in step 314; since the corpus can be used to create “a sentence classifier trained on the sentence corpus”, this means that the matrix is “from the machine learning model training”); and “analyzing edges in a network to determine significant association between nodes representing specific features” (Bhusan et al. FIG. 3A displays that the bacterial taxonomic abundance matrix is utilized to “construct a first bacterial association network” in Step 308 which is further used to construct a “feature count matrix” in Step 318 for applying a first machine learning model which is labeled as “a first classifier” and FIG. 3B also uses that same “feature count matrix” for generating a “refined association network” in Step 320 then applying a second machine learning model which is labeled as “a second classifier” in Step 322 ; [0084] further describes the “construct a first bacterial association network” in Step 308 as “A data normalization step is desirable in the next step to remove various sampling and experimental biases. The most common technique of normalization being a total sum scaling or a percentage normalization. However, advanced techniques including rarefaction based or centered log-ratio based transformation may also be used to normalize the abundance matrices. The bacterial taxonomic abundance matrix can then be used to identify relationship patterns between the bacteria (in the matrix) using proxy measures like significant correlations calculated from the matrix…Once the ‘all versus all’ pairwise bacterial associations (or in other words all possible unique pairwise combinations between the available bacteria) are calculated from the bacterial taxonomic abundance matrix, only the significant associations above or below a certain threshold value (which can be the correlation or the probability value of the association) are identified and an edge is connected between that pair of bacteria having a significant association (for example correlations having a probability value<0.05). Upon completion of the task for all the bacteria in the matrix, a bacterial association network is generated.”; “an edge is connected between that pair of bacteria” and “a bacterial association network” corresponds to “analyzing edges in a network”; “pair of bacteria having a significant association (for example correlations having a probability value < 0.05)” correspond to “determine significant associations between nodes representing specific features” in which each “bacteria” correspond to a “node” and “a probability value” corresponds to “specific features”) DeYoung et al. and Bhusan et al. are both analogous to the claimed invention because they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of DeYoung et al. to further incorporate generating a network from extracted features as taught by Bhusan et al. Doing so would extract bacterial associations from biomedical text more accurately (Bhusan et al. [0005]). Regarding claim 8, the rejections of claim 1 are incorporated. DeYoung et al. fails to teach: “resolving one or more specific features to compare to generic features by accessing additional resources to analyze specific features” However, Bhusan et al. teaches: “resolving one or more specific features to compare to generic features by accessing additional resources to analyze specific features” (Busan et al. FIG. 3A displays that the bacterial taxonomic abundance matrix is utilized to “construct a first bacterial association network” in Step 308 which is further used to construct a “feature count matrix” in Step 318 for applying a first machine learning model which is labeled as “a first classifier” and FIG. 3B also uses that same “feature count matrix” for generating a “refined association network” in Step 320 then applying a second machine learning model which is labeled as “a second classifier” in Step 322 ; [0088] describes how the bacterial association network is further used to construct the feature count matrix for machine learning prediction as “At step 316 of the method 300, the set of domain features is calculated for each of the abstracts present in the biomedical text corpus ‘Cz’ to generate a feature count matrix with one set of features for every abstract.”; [0117] “The results indicate that the set of domain features outperforms other methods using generic features like Sag of words' and 'TF-IDF'. All the measures including precision, recall and F1 score were best in all the four classifiers using the domain features namely Naïve Bayes, Logistic regression, Support Vector Machine (SVM) and Random forest classifier.”; “domain features is calculated” corresponds to “resolving one or more specific features”; “measures including precision recall and F1 score” correspond to “accessing additional resources to analyze specific features”) DeYoung et al. and Bhusan et al. are both analogous to the claimed invention because they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have replaced the ERASER corpora with specific features for analysis within literature as taught by DeYoung et al. with a bacterial abundance matrix with specific features of bacteria for analysis as taught by Bhusan et al. Hence, this would be a simple substitution of one known element (ERASER corpora) for another (bacterial abundance data) to obtain predictable results (analyzing specific features using additional metrics). (MPEP 2141(III)(B) Simple substitution of one known element for another to obtain predictable results). Claims 11-18 are system claims containing substantially the same elements as method claims 1-5 and 7-9, respectively. DeYoung et al. and Bhusan et al. teach the elements of claims 1-5 and 7-9, as shown above. Bhusan et al. also teaches: “a processor and a memory configured to provide computer program instructions to the processor to execute the function of code modules, the program instructions comprising” (Bhusan et al. [0006] “For example, in one embodiment, a system for annotation and classification of biomedical text having bacterial associations is provided. The system comprises a user interface, one or more hardware processors and a memory. The in communication with the one or more hardware processors, wherein the memory is in communication with one or more first hardware processors are configured to execute programmed instructions stored in the one or more first memories”) DeYoung et al. and Bhusan et al. are both analogous to the claimed invention because they are both in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have implemented the steps disclosed by DeYoung et al. on a processor and memory as disclosed by Bhusan et al. Therefore, this would be applying a known technique (applying evaluation metrics) to a known device (processor and memory) ready for improvement to yield predictable results (scoring machine learning models) (MPEP 2141(III)(D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results). Regarding claim 19, the rejections of claim 11 are incorporated. DeYoung et al. further mentions: “configuring analysis metrics for the ML models” (DeYoung et al. Section 6 “In Table 3 we evaluate models that perform discrete selection of rationales. We view these as inherently faithful, because by construction we know which snippets the decoder used to make a prediction.10 Therefore, for these methods we report only metrics that measure agreement with human annotations… We observe that for the “rationalizing” model of Lei et al. (2016), exploiting rationale-level supervision often (though not always) improves agreement with human-provided rationales”; “evaluate models” and “metrics that measure agreement with human annotation” corresponds to “analysis metrics for the machine learning models”) Claim 20 is a product claim containing substantially the same elements as system claim 11. DeYoung et al. and Bhusan et al. teach the elements of claim 11, as shown above. Bhusan et al. also teaches: “A computer program product for domain knowledge based evaluation of machine learning (ML) models for a target subject, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the instructions executable by a processor to cause the processor to” (Bhusan et al. [0006] “For example, in one embodiment, a system for annotation and classification of biomedical text having bacterial associations is provided. The system comprises a user interface, one or more hardware processors and a memory. The in communication with the one or more hardware processors, wherein the memory is in communication with one or more first hardware processors are configured to execute programmed instructions stored in the one or more first memories”; “one or more first memories” correspond to “computer readable storage medium”; “memory” and “processor” correspond to “A computer program product”) DeYoung et al. and Bhusan et al. are both analogous to the claimed invention because they are both in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have implemented the steps disclosed by DeYoung et al. on a processor and memory as disclosed by Bhusan et al. Therefore, this would be applying a known technique (applying evaluation metrics) to a known device (a computer program product) ready for improvement to yield predictable results (scoring machine learning models) (MPEP 2141(III)(D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.M.R./Examiner, Art Unit 2147 /ERIC NILSSON/Primary Examiner, Art Unit 2151
Read full office action

Prosecution Timeline

Dec 13, 2022
Application Filed
Oct 24, 2025
Non-Final Rejection mailed — §101, §102, §103
Jan 26, 2026
Response Filed
Aug 03, 2026
Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
4y 3m (~7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 2 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month