DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 6, 12-13, 17, 20-21, 26, and 28 are cancelled.
Claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are currently pending and under exam herein.
Claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are rejected.
Claims 10, 11, 14, 15, 24, and 25 are objected to.
Priority
The instant application claims benefit to PCT/IB2021/056870 filled July 28, 2021 and foreign application no. 10 200 000019180 from Italy filed Aug 4, 2020. The domestic and foreign priority benefit is acknowledged. Thus, the effective filing date of claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 is Aug 4, 2020.
Information Disclosure Statement
The information disclosure (IDS) was filed on 02/03/2023. All references in the IDS have been considered by the examiner and attached in this office action.
Drawings
The drawings filed on 02/03/2023 are accepted.
Specification
The spacing of the lines of the specification is such as to make reading difficult. New application papers with lines 1 1/2 or double spaced (see 37 CFR 1.52(b)(2)) on good quality paper are required.
The disclosure is objected to because it contains an embedded hyperlink and/or other form of browser-executable code.
Pg. 6: https://en.wikipedia.org/wiki/F1_score
Pg. 8: https://samtools.github.io/hts-specs/VCFv4.3.pdf
Pg. 13: https://evai.engenome.com
Pg. 15: http://clinvitae.invita.com
Applicant is required to delete the embedded hyperlink and/or other form of browser-executable code; references to websites should be limited to the top-level domain name without any prefix such as http:// or other browser-executable code. See MPEP § 608.01.
Claim Objections
Claim 10 is objected to because all acronyms need to be spell out within the claim. Specifically, the acronym “VCF” needs to be spell out.
Claims 11 and 15 are objected to because all acronyms need to be spell out within the claims. Specifically, the acronym “ACMG” in both claims needs to be spell out.
Claims 14, 24, and 25 are objected to because of the following informalities: “BS2, BS2” should read “BS1, BS2” instead in claims 14 and 24, and “BS2+BS2” should read “BS1 + BS2” in claim 25. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "the patient’s". There is insufficient antecedent basis for this limitation in the claim.
Claims 2-5, 7-11, 14-16, 18-19, 22-25, and 27 are also rejected under 35. USC 112(b) due to their dependency on claim 1.
Claim 5 recites that the trained algorithm is a Logistic Regression algorithm, but then also states that trained algorithm belongs to a group consisting of Decision Tree, Random Forest, Naïve Bayes, Gradient Boosting, and Support Vector Machine. However, Logistic Regression is a separate, unique algorithm that does not belong to any of the listed groups of other algorithms. Therefore, the claim is considered indefinite because there is a question or doubt as to what the trained algorithm encompasses. Applicant is asked to clarify what the exact metes and bounds of the trained algorithm are. For compact prosecution and examination purposes, Examiner will interpret that the trained algorithm is a Logistics Regression algorithm.
The terms “low frequency”, “low rate”, and “highly conserved” in claim 14 are relative terms which renders the claim indefinite. The terms “low frequency”, “low rate”, and “highly conserved” are not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
A broad range or limitation together with a narrow range or limitation that falls within the broad range or limitation (in the same claim) may be considered indefinite if the resulting claim does not clearly set forth the metes and bounds of the patent protection desired. See MPEP § 2173.05(c). In the present instance, claim 16 recites the broad recitation “wherein all of the criteria are used”, and the claim also recites “wherein a subset of the criteria is selected” which is the narrower statement of the range/limitation. The claim is considered indefinite because there is a question or doubt as to whether the feature introduced by such narrower language is (a) merely exemplary of the remainder of the claim, and therefore not required, or (b) a required feature of the claims.
Regarding claim 18, the phrase "i.e." renders the claim indefinite because it is unclear whether the limitations following the phrase are part of the claimed invention. See MPEP § 2173.05(d).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
In accordance with MPEP § 2106, claims found to recite statutory subject matter (Step 1: YES) are then analyzed to determine if the claims recite any concepts that equate to an abstract idea, law of nature or natural phenomenon (Step 2A, Prong 1). In the instant application, the claims recite the following limitations that equate to an abstract idea:
Claim 1 recites a method for determining the pathogenicity or benignity of a genomic variant in connection with a given disease (Abstract Idea: Mental Process and Natural Phenomenon). The overall concept of correlating a genetic variant with pathogenicity or benignity of a disease would constitute a natural phenomenon, as it is a naturally occurring correlation. In addition, this is a process that can be done within the human mind, as one can look at information regarding a variant and its association with a disease to determine if it is pathogenic (disease causing) or benign (not disease causing). The method starts by verifying for each variant whether or not the variant meets each of a plurality of predefined pathogenicity/benignity criteria (Abstract Idea: Mental Process). The pathogenicity/benignity criteria that are determine in this process are merely binary true or false propositions related to the variant, a first type condition (statistical condition and/or previous known condition) and a second type condition (condition specific of patient) (Abstract Idea: Mental Process). Again, this is a step that can be carried out in the human mind by looking at information regarding an established criteria associated with disease and information regarding the variant to determine if the variant meets the criteria before outputting a true or false statement. In addition, the method associates these criterion with levels of evidence, indicative of the level of pathogenicity or benignity (Abstract Idea: Mental Process). Which again, can be done in the human mind by looking at established information regarding the levels of evidence and the criteria. Next, the method prepares input information, which consists of information representing the number of pathogenicity/benignity criteria and level of evidence met by each variant (Abstract Idea: Mental Process). The act of associating a variant with numerous criteria and levels of evidence have already shown to be a mental process above. Therefore, merely putting the input information together, as recited, would also constitute a mental process. This input information is then process by a trained algorithm through machine learning techniques with training datasets (Abstract Idea: Mathematical Concept). A machine learning algorithm utilizes mathematical operations and correlations to find patterns within data, which in the broadest sense would make it a mathematical concept.
Claim 2 recites outputting an estimated probability of pathogenicity of at least one genomic variant (Abstract Idea: Mathematical Concept). This further proves that the algorithm is a mathematical process, that calculates and outputs an estimated probability based on known input data.
Claim 3 recites outputting a binary result representing whether or not the variant is pathogenic or benign, obtained by comparing the probability of pathogenicity with a respective threshold (Abstract Idea: Mental Process and/or Mathematical Concept). The process of looking at a probability and comparing it to a known threshold to see if it is greater (pathogenic) or less than (benign) is a mathematical operation that is well known. In addition, in the simplest sense, this operation can also be carried out in the human mind.
Claim 4 recites that the respective threshold is an optimized threshold that is determined based on the pre-training process (Abstract Idea: Mathematical Concept). The process of optimizing a threshold/variable, means running the algorithm multiple times in order to reduce an error function to come up with the most optimized threshold that is accurate to the data. The entire process is a mathematical operation and concept that is set to minimize error.
Claim 5 recites that the trained algorithm is a logistic regression algorithm (Abstract Idea: Mathematical Concept). A logistic regression algorithm itself is a mathematical concept that uses mathematical equations and operations to analyze data. In addition, claim 5 recites that the trained algorithm belongs to a group consisting of decision tree, random forest, naïve bayes, gradient boosting, support vector machine (Abstract Idea: Mathematical Concept). All of the above recited models utilize mathematical processing to analyze data and hence would be a mathematical concept.
Claim 8 recites that the training dataset in divided into three subsets, which includes the first and second subset, along with a third subset used as a test database (Abstract Idea: Mental Process). The claim specifies that the first subset is used as a training databases, the third subset is used as a test database to find the optimized threshold, and the second subset is used as a validation database to validate the optimized threshold. The process of dividing data into three subsets is something that can be practically done in the human mind, as one can organize data and divide it into subsets using mental evaluation and judgement. Therefore, these claims limitations constitute as a mental process. Claim 8 also recites calculating the precision and sensitivity of predictions at different decision thresholds to determine the optimal threshold (Abstract Idea: Mathematical Concept). Again, the process of optimizing a threshold/variable, means running the algorithm multiple times in order to reduce an error function to come up with the most optimized threshold that is accurate to the data. The precision and sensitivity are also variables that can be calculated and optimized to be at their maximum in order to ensure an accurate prediction based on known data. The entire process is a mathematical operation and concept that is set to minimize error.
Claim 16 recites that a subset of the criteria is selected based on the type of illness, and that all the criteria are used (Abstract Idea: Mental Process). The process of selecting criteria based on the type of illness consider is something that can be practically performed in the human mind, hence these limitations would constitute a mental process.
Claim 25 recites obtaining a sum corresponding to a group of corresponding criteria based on the equations listed below (Abstract Idea: Mathematical Process).
PNG
media_image1.png
188
352
media_image1.png
Greyscale
The act of summing up binary output (0’s and 1’s) to get a number is a mathematical concept.
Claim 27 recites modifying the input information for the trained algorithm by activating one or more pathogenicity/benignity criteria or changing the number of levels of evidence, or defining new criteria based on user input (Abstract Idea: Mental Process). The process of looking at user input and modifying information based on that user input (adding a criteria, changing the number of levels, adding new criteria) is a process that can be done in the human mind or with the assistance of a pen and paper.
These recitations are similar to the concepts of collecting information, analyzing it and displaying certain results of the collection and analysis in Electric Power Group, LLC, v. Alstom (830 F.3d 1350, 119 USPQ2d 1739 (Fed. Cir. 2016)), organizing and manipulating information through mathematical correlations in Digitech Image Techs., LLC v Electronics for Imaging, Inc. (758 F.3d 1344, 111 U.S.P.Q.2d 1717 (Fed. Cir. 2014)) and comparing information regarding a sample or test to a control or target data in Univ. of Utah Research Found. v. Ambry Genetics Corp. (774 F.3d 755, 113 U.S.P.Q.2d 1241 (Fed. Cir. 2014)) and Association for Molecular Pathology v. USPTO (689 F.3d 1303, 103 U.S.P.Q.2d 1681 (Fed. Cir. 2012)) that the courts have identified as concepts that can be practically performed in the human mind or mathematical relationships. Therefore, these limitations fall under the “Mental process” and “Mathematical concepts” groupings of abstract ideas. While claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 recite performing some aspects of the analysis with a “processor”, there are no additional limitations that indicate that this processor requires anything other than carrying out the recited mental process or mathematical concept in a generic computer environment. Merely reciting that a mental process is being performed in a generic computer environment does not preclude the steps from being performed practically in the human mind or with pen and paper as claimed. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then if falls within the “Mental processes” grouping of abstract ideas. As such, claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 recite an abstract idea (Step 2A, Prong 1: YES).
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further analyzed to determine if the claims as a whole integrate the recited judicial exception into a practical application or not (Step 2A, Prong 2). These judicial exceptions are not integrated into a practical application because the claims do not recite an additional element that reflects an improvement to technology or applies or uses the recited judicial exception in some other meaningful way. Rather, the instant claims recite additional elements that amount to mere instructions to implement the abstract idea in a generic computing environment or mere additional details regarding the data collection. Specifically, the claims recite the following additional elements:
Claim 1 recites accessing genomic data that is a list of a patient’s genomic variants, which amounts to mere data gathering under insignificant extra-solution activity. In addition, claim 1 recites outputting information from the trained algorithm that represents the pathogenicity/benignity of the genomic variants, which again is an insignificant extra-solution activity that equates to mere data outputting. Please see MPEP 2106.05(g) for specific details. Lastly, claim 1 recites a processor carries out the verifying if a variant meets the criteria and preparing the input data, which is equivalent to a generic computing environment.
Claim 2 and 3 again recites outputting data, which is insignificant extra-solution activity.
Claim 7 recites that there is a preliminary training step that utilizes two subsets of training dataset, a first subset used as a training database, and a second subset being used as a validation database. While this further specifies the processes in the trained algorithm and the type of data used to train the algorithm, it does not change the fact that the algorithm itself is a mathematical concept and the process of using the algorithm includes mathematical operations and comparisons.
Claim 9 recites that the first type condition is a statistical condition and/or a previously known condition that is verifiable on clinical databases while the second type condition is a patient specific condition verifiable on patient-specific input information. While these limitations further clarify what the first type and second type conditions are, they are only further specifying the input data into the algorithm. These limitations do not implement the judicial exception (the algorithm) into a practical use.
Claim 10 recites that the input genomic data is in a standard VCF format. Again, while this limitation further specifies the input data, it does not change the fact that the algorithm itself is a mathematical concept nor implement the algorithm into a practical use.
Claim 11 recites that the pathogenicity/benignity criteria are known criteria defined by ACMG (American College of Medical Genetics and Genomics) and are divided into subsets associated with various levels of evidence. While this claim limitations further specifies the input information for the algorithm, it does not change the fact that the algorithm itself is a mathematical concept.
Claim 14 further recites all the pathogenicity/benignity criteria with their established definition from the ACMG. While these definitions further specify the criteria used in the algorithm as an input data, it does not change the fact that the algorithm is still a mathematical concept.
Claim 15 adds another criteria, BP8, which is a non-ACMG criterion to the input information for the algorithm. Again, this is just further specifying the input data and does not change how the algorithm is fundamentally a mathematical concept.
Claim 18 recites the criteria related to the first type of condition (PVS1, PS1, PS3, PS4, PM1, PM2, PM4, PM5, PP2, PP3, PP5, BA1, BS1, BS3, BP1, BP3, BP4, BP7, BP8) and the criteria related to the second condition type (PS2, PM3, PM6, PP1, PP4, BS2, BS4, BP2, BP5). While this further defines the subsets of the criteria, it is merely limiting the type of input data, and does not change the fact that the algorithm is a mathematical concept.
Claim 19 recites that level of evidence is defined by known clinical standards and/or defined by ACMG. This limitation again further defines the subset of level of evidence associated with the criteria, but is merely limiting the type of input data.
Claim 22, 23, and 24 recites the levels of evidence, the criteria associated with each level, and how all the levels are used in the algorithm. Again, it is merely defining the input data into the algorithm and not implementing the judicial exception (the algorithm) is any meaningful way.
Claim 25 further recites the data structure of the input data into the algorithm, where each row is associated with a variant, each column is associated with a group of criteria, and each cell contains a number corresponding to the group of the criteria. Again, it is merely defining the input data into the algorithm and not implementing the judicial exception (the algorithm) is any meaningful way.
There are no limitations that indicate that the claimed processor or the formats of the provided data require anything other than generic computing systems. As such, these limitations equate to mere instructions to implement the abstract idea on a generic computer that the courts have stated does not render an abstract idea eligible in Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984. As such, claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are directed to an abstract idea (Step 2A, Prong 2: NO).
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims recite additional elements that equate to mere instructions to apply the recited exception in a generic computing environment and mere specifications on the input data that utilize established and known clinical criteria (from the ACMG). As discussed above, there are no additional limitations to indicate that the claimed processor requires anything other than generic computer components in order to carry out the recited abstract idea in the claims. Claims that amount to nothing more than an instruction to apply the abstract idea using a generic computer do not render an abstract idea eligible. Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984. In addition, the additional elements defining the type of input data for the trained algorithm are mere data gathering steps that are also based on known and established clinical standards for determining pathogenicity and benignity of variants (stated in claim 11, “wherein the pathogenicity/benignity criteria comprise criteria defined by known clinical standards and/or studies, and/or wherein the pathogenicity/benignity criteria comprise criteria defined by ACMG”). Furthermore, the inputting and/or outputting of data for analysis in a generic computer is a conventional and well-understood process that does not add significantly more to the judicial exceptions. The limitations of collecting information, analyzing it, and displaying certain results of the collection and analysis on a computer is merely linking the judicial exceptions to a particular technological environment, see Electric Power Group, LLC v. Alstom S.A., 830 F.3d 1350, 1354, 119 USPQ2d 1739, 1742 (Fed. Cir. 2016) for more details. Hence, the additional elements do not comprise an inventive concept when considered individually or as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. Therefore, the claims do not amount to significantly more than the judicial exception itself (Step 2B: No). As such, claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-5, 7-11, 14-16, 18-19, 22-25, and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Nicora et al. (Human Mutation Vol. 39 Issue 12 pgs. 1835-1846, Published Oct 9, 2018), herein after known as Nicora in view of Dong et al. (Human Molecular Genetics Vol. 24 No. 8 pgs. 2125-2137, Published Dec 30, 2015), herein after known as Dong. The claim limitations of the instant application are italicized below.
With respect to claim 1, Nicora teaches an automated system for guideline based variant classification in cardiovascular-related genes named CardioVAI (pg. 1835 Abstract, A method for determining the pathogenicity/benignity of a genomic variant in connection with a given disease). Nicora starts off by collecting two benchmark datasets of previously interpreted variants from online resources such as CardioDB and CLINVITAE (pg. 1838 left col para 1, accessing genomic data comprising a list of the patient’s genomic variants). Nicora then implements the 28 criteria proposed by the original ACMG-AMP guidelines, a BP8 criterion and any other criteria regarding patient specific information through manual inclusion (pg. 1837 left col para 1). Next, Nicora classifies the variants by checking if the variant meets the implement criteria, and assigns it a class depending on the number of criteria met and the corresponding level of evidence (pg. 1837 left col para 2, for each variant detected, verifying by a processor whether or not the variant meets each of a plurality of predefined pathogenicity/benignity criteria). Nicora specifically also looks at cardiovascular diseases (CVDs) for the application of CardioVAI, which encompass a broad set of disorders such as disease of myocardium, congenital heart disease, or disorders of the heart’s electrical circuit (pg. 1836 left col para 3, wherein said first type condition comprises a statistical condition and/or previous known condition, and said second type condition comprises a conditional specific of the patient). In addition, Nicora states that the criteria implementation incorporates CVDs specific knowledge gathered from omics-resources and CMP-EP guidelines, and that each gene-variant is associated to a list of CVDs conditions (pg. 1836 right col para 5, wherein each pathogenicity/benignity criterion is a proposition, which can be true or false, related to the variant, in connection with a first type condition or a second type condition, and wherein at least one of said pathogenicity/benignity criteria refers to a first type condition, and at least another one of the pathogenicity/benignity criteria refers to a second type condition). Furthermore, Nicora clarifies that each of the criteria used are associated with levels of evidence from the ACMG-AMP guidelines (pg. 1837 right col para 5 and Supp. Table S1, wherein each pathogenicity/benignity criterion is associated with a level of evidence, indicative of a condition or level of pathogenicity or benignity). Lastly, Nicora states that the level of evidence and pathogenicity/benignity criteria met by a variant are translated into scores, and the pathogenicity score for each variant-phenotype association is the sum of the scores of the triggered criteria (pg. 1837 right col para 5 and Supp. Table S1, for each variant and for each level of evidence, information representing the number of pathogenicity/benignity criteria associated with the level of evidence that are met by the variant).
Regarding claim 9, Nicora specifically looks at cardiovascular diseases (CVDs) for the application of CardioVAI, which encompass a broad set of disorders such as disease of myocardium, congenital heart disease, or disorders of the heart’s electrical circuit (pg. 1836 left col para 3). Nicora also clarifies that 17 of the 28 criteria proposed by the original ACMG-AMP guidelines have been automatically implemented in CardioVAI (pg. 1837 left col para 1). Namely PVS1, PS1, PS3, PS4, PM1, PM2, PM4, PM5, PP2, PP3, PP5, BA1, BS1, BS3, BP1, BP3, BP4, BP7, and BP8, are automatically implemented as established information regarding these criteria and their association with disease are already known (pg. 1837 left col para 1 and Supp. Table S1, wherein said first type condition comprises a statistical condition and/or a previous known condition which is verifiable on clinical or clinical-statistical databases accessible by the processor). While on the other hand, criteria that relied on patient specific information (such as co-segregation of the variant in family members) and patient specific diseases were not automatically included, but left for the patient to manually include (pg. 1837 left col para 1, said second type condition comprises a specific condition of the patient, which is verifiable based on patient-specific input information provided to the processor). Nicora elaborates that these criteria were PS2, PM3, PM6, PP1, PP4, BS2, BS4, BP2, BP5 (Supp. Table S1). In addition, Nicora states that the criteria implementation incorporates CVDs specific knowledge gathered from omics-resources and CMP-EP guidelines, and that each gene-variant is associated to a list of CVDs conditions (pg. 1836 right col para 5).
Concerning claim 10, Nicora states that the datasets were previously interpreted variants from online resources of CardioDB and CLINVITAE which provide genomic variant information in the form of variant call formats (pg. 1838 right col para 1, wherein the input genomic data are provided to the processor in a standard VCF format).
With respect to claim 11, Nicora states that the ACMG-AMP levels of evidence (very strong, stand alone, strong, moderate, and supporting) of each class (benign or pathogenic) were associated with each criteria (pg. 1837 right col para 5 and Supp. Table S1, wherein the pathogenicity/benignity criteria comprise pathogenicity criteria, the pathogenicity criteria being divided into subsets associated with various respective levels of evidence, and benignity criteria, the benignity criteria being divided into subsets associated with various respective levels of evidence, wherein the pathogenicity/benignity criteria comprise criteria defined by known clinical standards and/or studies, and/or wherein the pathogenicity/benignity criteria comprise criteria defined by ACMG).
Regarding claim 14 and claim 15, Nicora specifies the criteria used in their system, their definition, and their association with the levels of evidence in Supp. Table S1. The criteria are PVS1, PS1, PS2, PS3, PS4 PM1, PM2, PM3, PM4, PM5, PM6, PP1, PP2, PP3, PP4, PP5, BA1, BS2, BS2, BS3, BS4, BP1, BP2, BP3, BP4, BP5, BP6, BP7, and BP8 (Supp. Table S1, wherein the pathogenicity/benignity criteria comprise one or more of the following criteria: PVS1, PS1, PS2, PS3, PS4, PM1, PM2, PM3, PM4, PM5, PM6, PP1, PP2, PP3, PP4, PP5, BA1, BS2, BS2, BS3, BS4, BP1, BP2, BP3, BP4, BP5, BP6, BP7, wherein the pathogenicity/benignity criteria further comprise the following non-ACMG criterion: BP8: The same amino acid change has previously been determined to be benign, regardless of the type of nucleotide change). The definition of the criteria and their definitions are shown below in Supp. Table S1:
PNG
media_image2.png
894
710
media_image2.png
Greyscale
PNG
media_image3.png
892
715
media_image3.png
Greyscale
(Supp. Table S1, wherein said criteria are defined as follows: PVS1: Variant of the "null" type in a gene where the loss of function of the gene results in the onset of the disease is known; PS1: The same amino acid change has previously been interpreted as pathogenic, regardless of the type of nucleotide change; PS2: De novo variant confirmed in a patient with the disease and no family history (confirmed maternity and paternity); PS3: In vivo or in vitro functional studies confirm a damaging effect of the variant on the gene or gene product; PS4: The prevalence of the variant in individuals affected by the disease is significantly increased compared to the prevalence in controls; PM1: Variant located in a mutational hot-spot and/or in a critical and well-established functional domain, without benign variants; PM2: Variant absent in controls or at a very low frequency if the disease is recessive in Exome Sequencing Project, 1000 Genomes Project or Exome Aggregation Consortium;PM3: For recessive diseases, the variant is found in trans with a pathogenic variant; PM4: The protein length changes as a result of an in-frame deletion/insertion in a non- repeat region or stop-loss variants; PM5: Novel missense change at an amino acid residue where a different missense change was previously determined to be pathogenic; PM6: Presumed de novo variant, but without confirmation of paternity and maternity; PP1: Co-segregation with disease in multiple affected family members in a gene known to cause the disease; PP2: Missense variant in a gene which has a low rate of benign missense variants and in which missense-type variants cause the disease; PP3: Multiple evidences from computational tools support a deleterious effect of the variant on the gene or gene product; PP4: The patient's phenotype or family history is highly specific for the disease with a single genetic etiology; PP5: A reliable source reports the variant as pathogenic, but the evidence is not available to the laboratory to perform an independent assessment; BA1: The allele frequency of the variant is > 5% in Exome Sequencing Project, 1000 Genomes Project, or Exome Aggregation Consortium; BS1: The allele frequency is greater than that which would be expected for the disease; BS2: Variant observed in a healthy adult for a recessive (homozygous), dominant (heterozygous) or X-linked (hemizygous) disease, with full penetrance at a young age; BS3: In vivo or in vitro functional studies show no damaging effect of the variant on the gene or gene product; BS4: lack of segregation in affected family members; BP1: Missense variant in a gene for which primarily truncating variants are known to cause the disease; BP2: Observed in trans with a pathogenic variant for a dominant gene/disease and with full or observed penetrance in cis with a pathogenic variant in any inheritance pattern; BP3: In-frame deletion or insertion in a repetitive region without a known function; BP4: Multiple evidence from computational tools support a non-deleterious effect of the variant on the gene or gene product; BP5: Variant found in a case with an alternate molecular basis for the development of the disease; BP6: A reliable source reports the variant as benign, but the evidence is not available to the laboratory to perform an independent assessment; BP7: Synonymous (silent) variant for which the splicing prediction algorithms predict no impact on the splice sequence, nor the creation of a new splice site AND the nucleotide is highly conserved).
With respect to claim 16, Nicora demonstrates that the above 28 criteria are considered in the analysis of pathogenicity/benignity and some are selected based their correlation to cardio vascular diseases (pg. 1837 left col para 1 and Supp. Table S1, wherein a subset of the criteria is selected based on the type of illness or disease considered, wherein all of the criteria are used).
Regarding claim 18, Nicora clarifies that 17 of the 28 criteria proposed by the original ACMG-AMP guidelines have been automatically implemented in CardioVAI (pg. 1837 left col para 1). Namely PVS1, PS1, PS3, PS4, PM1, PM2, PM4, PM5, PP2, PP3, PP5, BA1, BS1, BS3, BP1, BP3, BP4, BP7, and BP8, are automatically implemented as established information regarding these criteria and their association with disease are already known (pg. 1837 left col para 1 and Supp. Table S1, wherein: the following criteria relate to a first type condition, i.e., to a statistical condition and/or a previous known condition: PVS1, PS1, PS3, PS4, PM1, PM2, PM4, PM5, PP2, PP3, PP5, BA1, BS1, BS3, BP1, BP3, BP4, BP7, BP8). While on the other hand, criteria that relied on patient specific information (such as co-segregation of the variant in family members) and patient specific diseases were not automatically included, but left for the patient to manually include (pg. 1837 left col para 1). Nicora elaborates that these criteria were PS2, PM3, PM6, PP1, PP4, BS2, BS4, BP2, BP5 (Supp. Table S1, and the following criteria relate to a second type condition, i.e., a condition specific of the patient: PS2, PM3, PM6, PP1, PP4, BS2, BS4, BP2, BP5).
Concerning claim 19, Nicora discloses that the levels of evidence correlated with the pathogenicity/benignity criteria are defined by the ACMG (pg. 1837 right col para 5, wherein the levels of evidence comprise levels of evidence associated with pathogenicity and levels of evidence associated with benignity, wherein the levels of evidence comprise levels defined by known clinical standards, and/or wherein the levels of evidence comprise levels of evidence defined by ACMG).
With respect to claim 22 and 23, Nicora discloses the levels of evidence regarding each class (pathogenic and benign) used as pathogenic: very strong, pathogenic: strong, pathogenic: moderate, pathogenic: supporting, benign: stand alone, benign: strong, benign: supporting (Supp. Table S1, see above for table, wherein the levels of evidence comprise one or more of the following levels of evidence: "Pathogenicity: Very Strong"; "Pathogenicity: Strong"; "Pathogenicity: Moderate"; "Pathogenicity: Supporting"; "Benignity: Stand Alone"; "Benignity: Very Strong"; "Benignity: Supporting", wherein all of the above levels of evidence are used).
Concerning claim 24, Nicora demonstrates how each criterion is associated with each level of evidence through the table, making each row of criterion align to each level of evidence that it is associated with (Supp. Table S1, see above for table, wherein the following associations apply: criterion PVS1 is associated with the level of evidence "Pathogenicity - Very Strong"; criteria PS1, PS2, PS3, PS4 are associated with the level of evidence "Pathogenicity - Strong"; criteria PM1, PM2, PM3, PM4, PM5, PM6 are associated with the level of evidence "Pathogenicity - Moderate"; criteria PP1, PP2, PP3, PP4, PP5 are associated with the level of evidence "Pathogenicity - Supporting"; criterion BA1 is associated with the level of evidence "Benignity - Stand Alone"; criteria BS2, BS2, BS3, BS4 are associated with the level of evidence "Benignity - Very Strong"; criteria BP1, BP2, BP3, BP4, BP5, BP6, BP7, BP8 are associated with the level of evidence "Benignity -Supporting").
With respect to claim 25, Nicora discloses that the pathogenicity score for each variant phenotype association is the sum of scores of the triggered criteria (pg. 1837 left col para 5 and pg. 1838 Table 4, each cell contains a number obtained from the sum corresponding to the group of the respective column, wherein each criterion of the group is associated with 1 if the criterion is met by the genetic variant of the respective row, and is associated with 0 if the criterion is not met by the genomic variant of the respective row; nPVS = PVS1, nPS = PS1 + PS2 + PS3 + PS4, npm = PM1 + PM2 + PM3 + PM4 + PM5 + PM6, nPP = PP1 + PP2 + PP3 + PP4 + PP5, nBA = BA1, nBS = BS1 + BS2 + BS3 + BS4, nBP = BP1 + BP2 + BP3 + BP4 + BP5 + PB6 + BP7 + BP8). In addition, Nicora discloses that the web-interface can output a table (CSV file) that shows each genetic variant as a row, with columns representing the class of variant (pathogenic/benign) and the criteria met by the variant (pg. 1843 Fig. 6, wherein said input information for the trained algorithm comprises, for each genomic variant, an indication of the number of pathogenicity/benignity criteria that are met by said genomic variant for each of the levels of evidence considered, wherein said input information for the trained algorithm comprises one or more tables, wherein: each row is associated with a respective genomic variant, each column is associated with a respective one of the following groups of criteria by level of evidence).
Regarding claim 27, Nicora discloses a web-interface that users can interact with and use to customize the ACMG-AMP rules for final classification by manually triggering or deactivating specific criteria and their level of evidence (pg. 1843 left col para 1 – right col para 1, modifying by a user through an electronic interface of said processor, the input information for the trained algorithm, before providing the input information as an input to the trained algorithm, wherein said modification step comprises activating one or more of said predefined pathogenicity/benignity criteria, and changing the number of the respective levels of evidence, or defining new criteria desired by the user and preparing the input information by inserting values related to said user- defined criteria).
However, Nicora does not explicitly disclose the use of a trained algorithm and processor (claims 1-5, 7-8, and 25) for predicting pathogenicity/benignity. Hence, Nicora does not teach the use of a logistic regression algorithm (claim 5), training data sets (claims 7-8) or optimizing a threshold through pretraining steps (claim 4). Therefore, Nicora also cannot teach obtaining output data from the algorithm and outputting the information (claims 1-3). Yet, the use of logistic regression algorithm to score and predict pathogenicity/benignity of genomic variants has been well established in the field of genomic analysis before the effective filing date of the instant application, as demonstrated by Dong below.
With respect to claim 1, Dong teaches the use of machine learning in the form of a logistic regression models to integrate multiple scoring methods for variant prediction (pg. 2126 left col para 2, processing said input information by the trained algorithm, wherein said trained algorithm is an algorithm trained by artificial intelligence techniques and/or machine learning techniques). Dong prepared the training data set by collecting variant information and performing feature selection and parameter tuning before inputting into the models (pg. 2126 right col para 1, preparing, by processing by the processor, input information for a trained algorithm). Dong discloses that the model was trained on a training set of 14191 deleterious/pathogenic mutations as true positive (TP) observations and 22001 neutral/benign mutations as true negative (TN) observations (pg. 2126 right col para 1, wherein said algorithm is trained in a preliminary training step, based on a training dataset of known cases, providing the algorithm to be trained with said input information calculated for each of the known cases, and training the algorithm based on the knowledge of the pathogenicity/benignity of the respective known cases). Dong also discloses that the logistic regression model utilizes scores from other variant classifier to create an ensemble score that represents if a variant is deleterious (pathogenic) or neutral (benign) (pg. 2130 left col para 3, obtaining output information from the trained algorithm, said output information representing the pathogenicity/benignity of each of the genomic variants considered).
Regarding claim 2, Dong discloses that the logistic regression model utilizes scores from other variant classifier to create an ensemble score that represents a probability of if a variant is deleterious (pathogenic) or neutral (benign) (pg. 2130 left col para 3, wherein said output information comprises an estimated probability of pathogenicity of at least one genomic variant considered, or of a plurality of genomic variants among the genomic variants considered, or of all the genomic variants considered).
Concerning claim 3, Dong discloses that they dichotomized their ensemble-based scores according to their standard dichotomizing threshold (0.5 for LR), with implications that scores above the threshold represented deleterious (pathogenic) variants and scores below the threshold represented neutral (benign) (pg. 2134 right col para 2, wherein the output information further comprises, for each genomic variant, a binary result representing whether the genomic variant is pathogenic or benign, wherein said binary result is obtained by comparing a probability of pathogenicity estimated for the genomic variant with a respective threshold, associated with the genomic variant).
With respect to claim 4, Dong discloses that for each ensemble score they evaluated, they varied the threshold for calling deleterious from the minimum value (0) to the maximum value (1) and computed the corresponding sensitivity and specificity in order to optimize the threshold (pg. 2135 left col para 5, wherein said respective threshold is an optimized threshold, common for all variants, and determined based on a pre-training).
Regarding claim 5, Dong teaches the use of machine learning in the form of a logistic regression models to integrate multiple scoring methods for variant prediction (pg. 2126 left col para 2, wherein said trained algorithm is a Logistic Regression algorithm).
Concerning claim 7, Dong first teaches an overall training set of 14191 deleterious/pathogenic mutations as true positive (TP) observations and 22001 neutral/benign mutations as true negative (TN) observations to train the logistic regression model (pg. 2126 right col para 1, a further preliminary training step, carried out based on two subsets of said training dataset containing data referring to known cases, a first subset being used as a training database). Then, Dong performs a 5-fold cross validation on the training dataset with the logistic regression model (pg. 2130 right col para 3, and a second subset being used as a validation database).
With respect to claim 8, Dong first teaches an overall training set of 14191 deleterious/pathogenic mutations as true positive (TP) observations and 22001 neutral/benign mutations as true negative (TN) observations that had three different parts in training the logistic regression model (pg. 2126 right col para 1, wherein said training dataset is divided into three subsets comprising, in addition to said first subset and second subset, also a third subset used as a test database). Dong discloses that the training dataset was used to initially train the logistic regression model (pg. 2126 right col para 1, and wherein the first subset is used as the training database). Then, Dong implies that the training data set was also used to optimize the threshold for calling deleterious mutations through calculations of sensitivity (True Positive/(True Positive + False Negative)) and specificity (1-False Positive (False Positive + True Negative) (pg. 2135 right col para 5, the third subset, is used to calculate precision and sensitivity of the prediction at different decision thresholds and to determine said optimized threshold, based on said calculation of precision and sensitivity at different thresholds). Although Dong does not explicitly disclose calculating/optimizing precision (True Positive/(True Positive + False Positive)), it would have been an obvious variable that can be calculated based on known rates to predict when the model makes a positive prediction, how often is it right. Lastly, Dong performs a 5-fold cross validation on the training dataset with the logistic regression model (pg. 2130 right col para 3, and the second subset is used as a validation database of the algorithm by setting said optimized threshold as a threshold).
It would have been prima facie obvious to one of ordinary skill in the art at the effective filing date of the invention to utilize the logistic regression model of Dong with the variant criteria data of Nicora to create an automatic and efficient system for clinical interpretation of sequence variants based on established clinical standards. One of ordinary skill in the art would have been motivated to integrate the machine learning framework of a logistic regression model to analyze the variant criteria data to improve statistical accuracy of prediction, scalability of data interpretation and efficiency of analysis. In addition, one of ordinary skill in the art before the effective filing date of the claimed invention would have a reasonable expectation of success at incorporating the logistic regression model of Dong with the variant data of Nicora as both operate on the same type of data (genomic variant data) in the same well-established field of variant interpretation. In addition, Dong has already shown success at integrating an assembly of different scores for pathogenic/benign variant prediction. Therefore, the application of a similar dataset within the well-establish frame work of logistic regression would have a predictable success rate.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENYU YANG whose telephone number is (571)272-0035. The examiner can normally be reached 8:30am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Olivia Wise can be reached at (571) 272-2249. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/W.Y./Examiner, Art Unit 1685
/OLIVIA M. WISE/Supervisory Patent Examiner, Art Unit 1685