DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-12 are currently pending and examined on the merits.
Priority
The instant application claims priority to U.S. Provisional Application 63/407,120 filed on 9/15/2022. At this point in examination, the effective filing date of claims 1-12 is 9/15/2022.
Information Disclosure Statement
No Information Disclosure Statement has been filed herein.
Specification
There are hyperlinks in pg. 7, para [0016], line 18 and pg. 8, para [0016], lines 1-2, 7, and 14. The disclosure is objected to because it contains an embedded hyperlink and/or other form of browser-executable code. Applicant is required to delete the embedded hyperlink and/or other form of browser-executable code; references to websites should be limited to the top-level domain name without any prefix such as http:// or other browser-executable code. See MPEP § 608.01.
Claim Objections
Claim 7 is objected to because of the following informalities:
In claim 7, line 24, there should be a colon “:” after “of”.
There is a typographical error. Appropriate correction is required.
Claim Interpretation - 35 USC § 112(f)
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f):
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
A: “a receiving module configured to receive an SNP profile derived from genome sequencing data of the subject” in claim 7.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
Para. [0020] of the published specification indicates that a processor may be implemented by a central processing unit (CPU), a microprocessor, a micro control unit (MCU), a system on a chip (SoC), or any circuit configurable/programmable in a software manner and/or hardware manner so as to implement the functionalities of the method. MPEP 2181.II.B. indicates that for computer-implemented means-plus-function limitations, the structure is an algorithm coupled with a microprocessor or computer. The above paragraph provides support for the processor or computer. The individual algorithms for each 112(f) invocation are detailed as follows:
A: “a receiving module” – para. [0019] of the instant specification states that the receiving module may be, but not limited to, a network interface controller or a wireless transceiver that supports wireless communication standards and is configured to receive SNP profiles transmitted by a remote electronic device (e.g., a computer), or a physical connector (e.g., a USB connector) that is configured to receive SNP profiles from an external electronic device (e.g., a flash drive) that is electronically connected to the receiving module. Therefore, the instant specification discloses a structure that performs the function of receiving SNP data. Thus, the description in the specification for the claimed receiving module has adequate corresponding structure.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f), applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite: (a) mathematical concepts, (e.g., mathematical relationships, formulas or equations, mathematical calculations); and (b) mental processes, i.e., concepts performed in the human mind, (e.g., observation, evaluation, judgement, opinion).
Subject matter eligibility evaluation in accordance with MPEP 2106:
Eligibility Step 1: Claims 1-6 are directed to a method (process) for evaluating a risk of a subject getting a specific disease. Claims 7-12 are directed to a system (machine). Therefore, these claims are encompassed by the categories of statutory subject matter, and thus satisfy the subject matter eligibility requirements under Step 1.
[Step 1: YES]
Eligibility Step 2A: First, it is determined in Prong One whether a claim recites a judicial exception, and if so, then it is determined in Prong Two whether the recited judicial exception is integrated into a practical application of that exception.
Eligibility Step 2A, Prong One: In determining whether a claim is directed to a judicial exception, examination is performed that analyzes whether the claim recites a judicial exception, i.e., whether a law of nature, natural phenomenon, or abstract idea is set forth described in the claim.
Claims 1-12 recite the following steps which fall within the mental processes and/or mathematical concepts groups of abstract ideas, as noted below.
Independent claim 1 further recites:
selecting, from an SNP profile derived from genome sequencing data of the subject, N number of target alleles that respectively match N number of specific risk alleles in the M number of specific risk alleles included in the reference database, N being a positive integer not greater than M (i.e., mental processes);
selecting, from among the M number of original parameter sets, N number of target parameter sets that correspond respectively to the N number of specific risk alleles (i.e., mental processes);
calculating, for each of the N number of target parameter sets, a race factor based on the global risk allele frequency and the group-specific risk allele frequency of the target parameter set (i.e., mental processes, mathematical concepts);
calculating a genetic factor based on the statistics respectively of the N number of target parameter sets, the global reference allele frequencies respectively of the N number of target parameter sets, the race factors respectively calculated for the N number of target parameter sets, and the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets (i.e., mental processes, mathematical concepts);
calculating a citation factor based on the numbers of citation times respectively of the N number of target parameter sets (i.e., mental processes, mathematical concepts);
calculating a risk score based on the genetic factor and the citation factor (i.e., mental processes, mathematical concepts).
Dependent claims 2 and 8 further recite:
wherein for each of the M number of original parameter sets, the statistics include: a p-value representing a probability that an association of the specific disease with the corresponding one of the M number of specific risk alleles is due to random chance; and an odds ratio, which is a ratio of a probability of a person with the corresponding one of the M number of specific risk alleles getting the specific disease to a probability of a person without the corresponding one of the M number of specific risk alleles getting the specific disease (i.e., mental processes, mathematical concepts; this is further information limiting the judicial exceptions).
Dependent claims 3 and 9 further recite:
wherein the step of calculating a genetic factor is to calculate the genetic factor according to a formula:
F
a
c
t
o
r
G
e
n
e
t
i
c
=
1
M
∑
i
=
1
N
-
log
10
P
i
×
O
R
i
×
S
N
P
_
T
y
p
e
i
×
F
a
c
t
o
r
R
a
c
e
,
i
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
e
f
,
i
, where
F
a
c
t
o
r
G
e
n
e
t
i
c
represents the genetic factor,
P
i
represents the p-value for an
i
t
h
one of the N number of specific risk alleles,
O
R
i
represents the odds ratio for the
i
t
h
one of the N number of specific risk alleles,
S
N
P
_
T
y
p
e
i
represents the number of chromosomes in a homologous chromosome pair having the
i
t
h
one of the N number of specific risk alleles,
F
a
c
t
o
r
R
a
c
e
,
i
represents the race factor for the
i
t
h
one of the N number of specific risk alleles, and
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
e
f
,
i
represents the global reference allele frequency for the
i
t
h
one of the N number of specific risk alleles (i.e., mental processes, mathematical concepts).
Dependent claims 4 and 10 further recite:
wherein for an
i
t
h
one of the target parameter sets that corresponds to an
i
t
h
one of the N number of specific risk alleles, i being an integer ranging from one to N, the step of calculating a race factor is to calculate the race factor according to formulas:
F
a
c
t
o
r
R
a
c
e
,
i
=
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
+
1
,
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
≥
0
1
1
-
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
,
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
<
0
; and
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
=
F
r
e
q
u
e
n
c
y
G
r
o
u
p
r
i
s
k
,
i
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
i
s
k
,
i
, where
F
a
c
t
o
r
R
a
c
e
,
i
represents the race factor for the
i
t
h
one of the N number of specific risk alleles,
F
r
e
q
u
e
n
c
y
G
r
o
u
p
r
i
s
k
,
i
represents the group-specific risk allele frequency for the
i
t
h
one of the N number of specific risk alleles, and
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
i
s
k
,
i
represents the global risk allele frequency for the
i
t
h
one of the N number of specific risk alleles (i.e., mental processes, mathematical concepts).
Dependent claims 5 and 11 further recite:
wherein the step of calculating a citation factor is to calculate the citation factor according to a formula:
F
a
c
t
o
r
c
i
t
a
t
i
o
n
=
ln
∑
i
=
1
N
(
C
i
t
a
t
i
o
n
n
u
m
i
+
1
)
, where
F
a
c
t
o
r
c
i
t
a
t
i
o
n
represents the citation factor, and
C
i
t
a
t
i
o
n
n
u
m
i
represents the number of citation times for an
i
t
h
one of the N number of specific risk alleles (i.e., mental processes, mathematical concepts).
Dependent claims 6 and 12 further recite:
wherein the step of calculating a risk score is to calculate the risk score according to a formula:
S
c
o
r
e
r
i
s
k
=
100
,
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
>
100
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
,
0
<
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
≤
100
, where
S
c
o
r
e
r
i
s
k
represents the risk score,
F
a
c
t
o
r
G
e
n
e
t
i
c
represents the genetic factor, and
F
a
c
t
o
r
C
i
t
a
t
i
o
n
represents the citation factor (i.e., mental processes, mathematical concepts).
Independent claim 7 further recites:
selecting, from the SNP profile derived from genome sequencing data of the subject, N number of target alleles that respectively match N number of specific risk alleles in the M number of specific risk alleles indicated in the reference database, N being a positive integer not greater than M (i.e., mental processes);
selecting, from among the M number of original parameter sets, N number of target parameter sets that correspond respectively to the N number of specific risk alleles (i.e., mental processes);
calculating, for each of the N number of target parameter sets, a race factor based on the global risk allele frequency and the group-specific risk allele frequency of the target parameter set (i.e., mental processes, mathematical concepts);
calculating a genetic factor based on the statistics respectively of the N number of target parameter sets, the global reference allele frequencies respectively of the N number of target parameter sets, the race factors respectively calculated for the N number of target parameter sets, and the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets (i.e., mental processes, mathematical concepts);
calculating a citation factor based on the numbers of citation times respectively of the N number of target parameter sets (i.e., mental processes, mathematical concepts);
calculating a risk score based on the genetic factor and the citation factor (i.e., mental processes, mathematical concepts).
The abstract ideas recited in the claims are evaluated under the broadest reasonable interpretation (BRI) of the claim limitations when read in light of and consistent with the specification. As the claims are currently recited, the method could be performed by writing down the calculations for the race factors, genetic factors, citation factors, and risk scores with pen and paper. Therefore, writing out these calculations can be practically performed in the human mind. Additionally, the recited limitations that are identified as judicial exceptions from the mathematical concepts grouping of abstract ideas are abstract ideas irrespective of whether or not the limitations are practical to perform in the human mind.
Therefore, claims 1-12 recite an abstract idea.
[Step 2A, Prong One: YES]
Eligibility Step 2A, Prong Two: In determining whether a claim is directed to a judicial exception, further examination is performed that analyzes if the claim recites additional elements that, when examined as a whole, integrates the judicial exception(s) into a practical application (MPEP 2106.04(d)). A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. The claimed additional elements are analyzed to determine if the abstract idea is integrated into a practical application (MPEP 2106.04(d)(I); MPEP 2106.05(a-h)). If the claim contains no additional elements beyond the abstract idea, the claim fails to integrate the abstract idea into a practical application (MPEP 2106.04(d)(III)).
The judicial exceptions identified in Eligibility Step 2A, Prong One are not integrated into a practical application because of the reasons noted below.
Claims 1 and 7 recite the additional non-abstract elements of data gathering:
storing a reference database by collecting data from a medical literature database, an allele frequency database, and a plurality of databases that compiles data of genome-wide association study (GWAS), the reference database containing M number of original parameter sets that respectively correspond to M number of specific risk alleles respectively at M number of chromosomal positions where single-nucleotide polymorphisms (SNPs) related to the specific disease occur, M being a positive integer greater than one, each of the M number of original parameter sets including a plurality of statistics related to the corresponding one of the M number of specific risk alleles, a global risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in global population, a group-specific risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in a certain race group, a global reference allele frequency that is related to the global risk allele frequency, a number of citation times that literatures related to the corresponding one of the M number of specific risk alleles are cited, and a number of chromosomes in a homologous chromosome pair having the corresponding one of the M number of specific risk alleles (claim 1);
a receiving module configured to receive an SNP profile derived from genome sequencing data of the subject (claim 7).
Data gathering steps are not an abstract idea, they are extra-solution activity, as they collect the data needed to carry out the JE. The data gathering does not impose any meaningful limitation on the JE, or how the JE is performed. The additional limitation (data gathering) must have more than a nominal or insignificant relationship to the identified judicial exception. (MPEP 2106.04/.05, citing Intellectual Ventures LLC v. Symantee Corp, McRO, TLI communications, OIP Techs. Inc. v. Amason.com Inc., Electric Power Group LLC v. Alstrom S.A.).
Claim 7 recites the additional non-abstract element (EIA) of a general-purpose computer system or parts thereof:
a system comprising a processor (claim 7);
a storage configured to store a reference database that is established in advance by collecting data from a medical literature database, an allele frequency database, and a plurality of databases that compiles data of genome-wide association study (GWAS), the reference database containing M number of original parameter sets that respectively correspond to M number of specific risk alleles respectively at M number of chromosomal positions where single-nucleotide polymorphisms (SNPs) related to the specific disease occur, M being a positive integer greater than one, each of the M number of original parameter sets including a plurality of statistics related to the corresponding one of the M number of specific risk alleles, a global risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in global population, a group-specific risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in a certain race group, a global reference allele frequency that is related to the global risk allele frequency, a number of citation times that literatures related to the corresponding one of the M number of specific risk alleles are cited, and a number of chromosomes in a homologous chromosome pair having the corresponding one of the M number of specific risk alleles (claim 7).
The EIA do not provide any details of how specific structures of the computer elements are used to implement the JE. The claims require nothing more than a general-purpose computer to perform the functions that constitute the judicial exceptions. The computer elements of the claims do not provide improvements to the functioning of the computer itself (as in DDR Holdings, LLC v. Hotels.com LP); they do not provide improvements to any other technology or technical field (as in Diamond v. Diehr); nor do they utilize a particular machine (as in Eibel Process Co. v. Minn. & Ont. Paper Co.). Hence, these are mere instructions to apply the JE using a computer, and therefore the claim does not recite integrate that JE into a practical application.
Thus, the additionally recited elements merely invoke a computer as a tool, and/or amount to insignificant extra-solution data gathering activity, and as such, when all limitations in claims 1-12 have been considered as a whole, the claims are deemed to not recite any additional elements that would integrate a judicial exception into a practical application. Claims 1 and 7 contain additional elements that would not integrate a judicial exception into a practical application and are further probed for inventive concept in Step 2B.
[Step 2A, Prong Two: NO]
Eligibility Step 2B: Because the claims recite an abstract idea, and do not integrate that abstract idea into a practical application, the claims are probed for a specific inventive concept. The judicial exception alone cannot provide that inventive concept or practical application (MPEP 2106.05). Identifying whether the additional elements beyond the abstract idea amount to such an inventive concept requires considering the additional elements individually and in combination to determine if they amount to significantly more than the judicial exception (MPEP 2106.05A i-vi).
The claims do not include any additional elements that are sufficient to amount to significantly more than the judicial exception(s) because of the reasons noted below.
With respect to claims 1 and 7: The limitations identified above as non-abstract elements (EIA) related to data gathering do not rise to the level of significantly more than the judicial exception. Activities such as data gathering do not improve the functioning of a computer, or comprise an improvement to any other technical field. The limitations do not require or set forth a particular machine, they do not affect a transformation of matter, nor do they provide an unconventional step (citing McRO and Trading Technologies Int’l v. IBG). Data gathering steps constitute a general link to a technological environment. Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception are insufficient to provide significantly more (as discussed in Alice Corp.,).
With respect to claim 7: The limitations identified above as non-abstract elements (EIA) related to general-purpose computer systems do not rise to the level of significantly more than the judicial exception. These elements do not improve the functioning of the computer itself, or comprise an improvement to any other technical field (Trading Technologies Int’l v. IBG, TLI Communications). They do not require or set forth a particular machine (Ultramercial v. Hulu, LLC., Alice Corp. Pty. Ltd v. CLS Bank Int’l), they do not affect a transformation of matter, nor do they provide an unconventional step. Simply appending well understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception are insufficient to provide significantly more (as discussed in Alice Corp., CyberSource v. Retail Decisions, Parker v. Flook, Versata Development Group v. SAP America).
[Step 2B: NO]
Therefore, claims 1-12 are patent ineligible under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2 are rejected under 35 U.S.C. 103 as being unpatentable over Polfus et al. (Human Genetics and Genomics Advances, 2021, 2(2), 1-42), in view of Chunn et al. (Frontiers in Genetics, 2020, 11, 1-22) and Folkersen et al. (Frontiers in Genetics, 2020, 11, 1-11).
With respect to claim 1:
Regarding the recited storing a reference database by collecting data from a medical literature database, allele frequency database, and a plurality of databases that compiles data of genome-wide association study (GWAS), Polfus et al. discloses performing a multiethnic GWAS meta-analysis of type 2 diabetes (T2D), combining population-specific GWAS results for T2D in PAGE with European ancestry GWAS results from the Diabetes Genetics Replication and Meta-analysis (DIAGRAM) consortium to discover risk loci for T2D (pg. 2, col. 1, para. 2, lines 4-9). Table S2 depicts multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database of allele frequencies built with GWAS data.
Polfus et al. does not disclose a medical literature database.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a medical literature database.
Chunn et al. does not disclose the reference database containing M number of original parameter sets that respectively correspond to M number of specific risk alleles respectively at M number of chromosomal positions where single-nucleotide polymorphisms (SNPs) related to the specific disease occur, M being a positive integer greater than one, each of the M number of original parameter sets including a plurality of statistics related to the corresponding one of the M number of specific risk alleles.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). Also, further discloses that this data was derived from population-specific and multiethnic meta-analyses conducted across PAGE studies and DIAGRAM for 28 million SNPs (pg. 3, col. 2, para. 4, lines 1-3). This teaches a database comprising parameter sets corresponding to specific risk alleles respectively at different chromosomal positions where SNPs related to type 2 diabetes occur. Each of the parameter sets include summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Chunn et al. does not disclose a global risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in global population.
However, Polfus et al. discloses a risk allele frequency corresponding to a risk allele at chromosome position rs13052926 for multiethnic ancestry group (pg. 5, Table 1; pg. 5, col. 2, para. 2, lines 1-3). This teaches global risk allele frequency corresponding to a variant that is common in all ancestry groups combined.
Chunn et al. does not disclose a group-specific risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in a certain race group.
However, Polfus et al. discloses risk allele frequencies corresponding to specific risk alleles for each ancestry group (pg. 5, Table 1). This teaches group-specific risk allele frequencies corresponding to specific risk alleles in certain race groups.
Chunn et al. does not disclose a global reference allele frequency that is related to the global risk allele frequency.
However, Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. does not disclose a number of citation times that literatures related to the corresponding one of the M number of specific risk alleles are cited.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citations that literature related to specific variants are cited.
Polfus et al. and Chunn et al. do not disclose a number of chromosomes in a homologous chromosome pair having the corresponding one of the M number of specific risk alleles.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
Chunn et al. and Folkersen et al. do not disclose selecting, from an SNP profile derived from genome sequencing data of the subject, N number of target alleles that respectively match N number of specific risk alleles in the M number of specific risk alleles included in the reference database, N being a positive integer not greater than M.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that population-specific and multiethnic meta-analyses were conducted across PAGE studies and DIAGRAM for 28 million SNPs (pg. 3, col. 2, para. 4, lines 1-3). This teaches selecting target risk loci that match the specific risk alleles included in the reference database.
Chunn et al. and Folkersen et al. do not disclose selecting, from among the M number of original parameter sets, N number of target parameter sets that correspond respectively to the N number of specific risk alleles.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). This teaches selecting parameter sets of specific risk loci that are statistically significant.
Chunn et al. and Folkersen et al. do not disclose calculating, for each of the N number of target parameter sets, a race factor based on the global risk allele frequency and the group-specific risk allele frequency of the target parameter set.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Chunn et al. does not disclose calculating a genetic factor based on the statistics respectively of the N number of target parameter sets, the global reference allele frequencies respectively of the N number of target parameter sets, the race factors respectively calculated for the N number of target parameter sets, and the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database comprising parameter sets including summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Polfus et al. does not disclose the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
It would be obvious to calculate a genetic factor based on the elements disclosed by Polfus et al. and Folkersen et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk.
Polfus et al. and Folkersen et al. do not disclose calculating a citation factor based on the numbers of citation times respectively of the N number of target parameter sets.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citation times. It would be obvious to calculate a citation factor based on citation times because Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Therefore, a citation factor based on citation times provides reliability in variant interpretation.
Chunn et al. does not disclose calculating a risk score based on the genetic factor and the citation factor.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database comprising parameter sets including summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Polfus et al. does not disclose the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
It would be obvious to calculate a genetic factor based on the elements disclosed by Polfus et al. and Folkersen et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk.
Polfus et al. and Folkersen et al. do not disclose a citation factor.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citation times. It would be obvious to calculate a citation factor based on citation times because Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Therefore, a citation factor based on citation times provides reliability in variant interpretation.
It would be obvious to calculate a risk score based on the genetic factor disclosed by Polfus et al. and Folkersen et al. and the citation factor disclosed by Chunn et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk. Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Also, further discloses that citation counts is a measure of relevance (pg. 5, col. 2, para. 1, lines 1-6). Therefore, a citation factor provides reliability in variant interpretation and can help determine impact of citations on variants associated with disease risk.
With respect to claim 2:
Chunn et al. and Folkersen et al. do not disclose wherein for each of the M number of original parameter sets, the statistics include: a p-value representing a probability that an association of the specific disease with the corresponding one of the M number of specific risk alleles is due to random chance.
However, Polfus et al. discloses observing statistically significant associations in 39 of the 582 known risk loci in the multiethnic meta-analysis, which includes rs35011184 near TCF7L2 (MIM: 602228) (OR = 1.31, 95% confidence interval [CI] = 1.27–1.34, p = 3.32
×
10
-
102
), which was also the most significant variant in European (OR = 1.34, 95% CI = 1.29–1.38, p = 1.55
×
10
-
75
), and African populations (OR = 1.29, 95% CI = 1.20–1.39, p = 2.42
×
10
-
11
) (pg. 4, col. 2, para. 1, lines 7-25). Also, further discloses that summary statistics from each contributing GWAS were combined across studies to form a single combined log odds ratio (OR) estimate, standard error, and Wald test for each variant (pg. 3, col. 2, para. 4, lines 3-7). This teaches p-values as probabilities that risk loci are associated with type 2 diabetes. The lower the p-value, the more statistically significant the associations are.
Chunn et al. and Folkersen et al. do not disclose an odds ratio, which is a ratio of a probability of a person with the corresponding one of the M number of specific risk alleles getting the specific disease to a probability of a person without the corresponding one of the M number of specific risk alleles getting the specific disease.
However, Polfus et al. discloses odds ratios calculated for each risk/other allele in Table 1 (pg. 5, Table 1). Also, further discloses that summary statistics from each contributing GWAS were combined across studies to form a single combined log odds ratio (OR) estimate, standard error, and Wald test for each variant (pg. 3, col. 2, para. 4, lines 3-7). This teaches odds ratios as a ratio of a person with risk alleles getting type 2 diabetes to the person without the risk allele getting type 2 diabetes.
It would have been prima facie obvious to one of ordinary skill in the art to modify the multiethnic GWAS study disclosed by Polfus et al. to incorporate citation counts disclosed by Chunn et al. and number of chromosomes disclosed by Folkersen et al. One would be motivated to incorporate citation counts and chromosome counts because Chunn et al. discloses that Mastermind demonstrated a sensitivity of 98.4% compared to 4.4, 45.6, and 37.4% for alternatives PubMed, Google Scholar, and ClinVar (pg. 1, Abstract, lines 15-17). Therefore, incorporating citation counts from Mastermind in the multiethnic GWAS study is very reliable. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the multiethnic GWAS study helps advance public perception of genomic risk. There is a likelihood of success, since all methods are of genetic variant analyses and disease risk scoring, which are well known techniques in the field of bioinformatics.
Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Polfus et al. (Human Genetics and Genomics Advances, 2021, 2(2), 1-42), in view of Chunn et al. (Frontiers in Genetics, 2020, 11, 1-22), Folkersen et al. (Frontiers in Genetics, 2020, 11, 1-11), and Sadat et al. (Texas Instruments, 2015, 1-8).
With respect to claim 7:
Claim 7 recites a system comprising a storage and a processor.
Polfus et al. discloses example code for the GRS analysis performed in the study as well as code for analyses using other indicated software being readily available from the websites of the corresponding software (pg. 8, col. 2, para. 1). The example code requires an operating system of a general-purpose computer such as Windows, Mac, or LINUX. These general-purpose computers require the recited system comprising a storage and a processor.
Regarding the recited collecting data from a medical literature database, allele frequency database, and a plurality of databases that compiles data of genome-wide association study (GWAS), Polfus et al. discloses performing a multiethnic GWAS meta-analysis of type 2 diabetes (T2D), combining population-specific GWAS results for T2D in PAGE with European ancestry GWAS results from the Diabetes Genetics Replication and Meta-analysis (DIAGRAM) consortium to discover risk loci for T2D (pg. 2, col. 1, para. 2, lines 4-9). Table S2 depicts multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database of allele frequencies built with GWAS data.
Polfus et al. does not disclose a medical literature database.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a medical literature database.
Chunn et al. does not disclose the reference database containing M number of original parameter sets that respectively correspond to M number of specific risk alleles respectively at M number of chromosomal positions where single-nucleotide polymorphisms (SNPs) related to the specific disease occur, M being a positive integer greater than one, each of the M number of original parameter sets including a plurality of statistics related to the corresponding one of the M number of specific risk alleles.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). Also, further discloses that this data was derived from population-specific and multiethnic meta-analyses conducted across PAGE studies and DIAGRAM for 28 million SNPs (pg. 3, col. 2, para. 4, lines 1-3). This teaches a database comprising parameter sets corresponding to specific risk alleles respectively at different chromosomal positions where SNPs related to type 2 diabetes occur. Each of the parameter sets include summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Chunn et al. does not disclose a global risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in global population.
However, Polfus et al. discloses a risk allele frequency corresponding to a risk allele at chromosome position rs13052926 for multiethnic ancestry group (pg. 5, Table 1; pg. 5, col. 2, para. 2, lines 1-3). This teaches global risk allele frequency corresponding to a variant that is common in all ancestry groups combined.
Chunn et al. does not disclose a group-specific risk allele frequency that is related to an allele frequency of the corresponding one of the M number of specific risk alleles in a certain race group.
However, Polfus et al. discloses risk allele frequencies corresponding to specific risk alleles for each ancestry group (pg. 5, Table 1). This teaches group-specific risk allele frequencies corresponding to specific risk alleles in certain race groups.
Chunn et al. does not disclose a global reference allele frequency that is related to the global risk allele frequency.
However, Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. does not disclose a number of citation times that literatures related to the corresponding one of the M number of specific risk alleles are cited.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citations that literature related to specific variants are cited.
Polfus et al. and Chunn et al. do not disclose a number of chromosomes in a homologous chromosome pair having the corresponding one of the M number of specific risk alleles.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
Polfus et al., Chunn et al., and Folkersen et al. do not disclose a receiving module configured to receive an SNP profile derived from genome sequencing data of the subject.
However, Sadat et al. discloses a laptop having a USB connector used to connect USB peripherals such as a hard disk and flash drive for data transfer (pg. 3, col. 1, para. 2, lines 1-9). This teaches a USB connector as a receiving module that sends and receives data between a computer and an external drive.
Chunn et al., Folkersen et al., and Sadat et al. do not disclose selecting, from an SNP profile derived from genome sequencing data of the subject, N number of target alleles that respectively match N number of specific risk alleles in the M number of specific risk alleles included in the reference database, N being a positive integer not greater than M.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that population-specific and multiethnic meta-analyses were conducted across PAGE studies and DIAGRAM for 28 million SNPs (pg. 3, col. 2, para. 4, lines 1-3). This teaches selecting target risk loci that match the specific risk alleles included in the reference database.
Chunn et al., Folkersen et al., and Sadat et al. do not disclose do not disclose selecting, from among the M number of original parameter sets, N number of target parameter sets that correspond respectively to the N number of specific risk alleles.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). This teaches selecting parameter sets of specific risk loci that are statistically significant.
Chunn et al., Folkersen et al., and Sadat et al. do not disclose calculating, for each of the N number of target parameter sets, a race factor based on the global risk allele frequency and the group-specific risk allele frequency of the target parameter set.
However, Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Chunn et al. and Sadat et al. do not disclose calculating a genetic factor based on the statistics respectively of the N number of target parameter sets, the global reference allele frequencies respectively of the N number of target parameter sets, the race factors respectively calculated for the N number of target parameter sets, and the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database comprising parameter sets including summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Polfus et al. does not disclose the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
It would be obvious to calculate a genetic factor based on the elements disclosed by Polfus et al. and Folkersen et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk.
Polfus et al., Folkersen et al., and Sadat et al. do not disclose calculating a citation factor based on the numbers of citation times respectively of the N number of target parameter sets.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citation times. It would be obvious to calculate a citation factor based on citation times because Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Therefore, a citation factor based on citation times provides reliability in variant interpretation.
Chunn et al. and Sadat et al. do not disclose calculating a risk score based on the genetic factor and the citation factor.
However, Polfus et al. discloses Table S2 depicting multiethnic and population-specific effects for each risk allele in each T2D loci, including risk allele frequencies, p-values, and odds ratios per ethnic group (pg. 3-10, Supplementary Data, Table S2). This teaches a database comprising parameter sets including summary statistics such as p-values and odds ratios corresponding to the specific risk alleles.
Polfus et al. discloses imputing allele dosages for autosomal variants in PAGE using the 1000 Genomes phase 3 reference panel (pg. 3, col. 1, para. 4). Also, further discloses that recent development of genetic imputation reference panels and genotyping arrays that capture common variation in ancestrally diverse groups allows for better characterization of the genetic architecture of T2D across race and ethnicity (pg. 2, col. 1, para. 1, lines 10-14). This teaches that the 1000 Genomes reference panel supplies the global reference allele frequencies for better characterization of type 2 diabetes across ethnicity groups.
Polfus et al. discloses observing genome-wide significant associations for 39 of the 582 known risk loci in the multiethnic meta-analysis depicted in Table S2, 20 of which were also significant in population-specific analyses (pg. 4, col. 2, para. 1, lines 7-25; pg. 3, Supplementary Data, Table S2). Also, further discloses that variant rs13052926 is common in all ancestry groups, and a genome-wide significant association for this variant was observed in the multiethnic meta-analysis with the strongest associations observed in the European population, followed by the Asian, Hispanic, and African populations, with no evidence of association observed in Native Hawaiians (pg. 5-6, col. 2, para. 2; pg. 5, Table 1). Polfus et al. discloses that the findings from the meta-analyses underscore the importance of genetic discovery and risk characterization in diverse populations and the urgent need to further increase representation of non-European ancestry individuals in genetics research to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). This teaches considering the different ethnicity groups while making observations in the overall multiethnic population combined. It would be obvious to calculate a race factor based on global risk allele frequency and group-specific risk allele frequency because these observations help improve genetic-based risk prediction of different ethnicity groups within a global population.
Polfus et al. does not disclose the numbers of chromosomes in homologous chromosome pairs of the N number of target parameter sets.
However, Folkersen et al. discloses a top-SNP calculation mode for polygenic score calculation that incorporates effect allele counts from genotype data (0, 1, or 2) (pg. 9, col. 1, para. 3). An individual has two homologous chromosomes at each locus. Therefore, alleles can be on 0, 1, or both chromosomes. This teaches the number of chromosomes in a pair of chromosomes having a corresponding risk allele.
It would be obvious to calculate a genetic factor based on the elements disclosed by Polfus et al. and Folkersen et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk.
Polfus et al., Folkersen et al., and Sadat et al. do not disclose a citation factor.
However, Chunn et al. discloses a genomic search engine Mastermind that allows users to search a comprehensive dataset comprising medical literature, pre-annotated for genetic content, and determine relevance of papers by the presence of genes or variants in the title or citation count (pg. 2, col. 2, para. 2, lines 1-4; pg. 5, col. 2, para. 1, lines 1-6). Also, further discloses the Cited Variants Reference (CVR), which is a useful tool for evaluation of genomic associations and is a download of Mastermind’s database that includes all variants affecting less than 4 nucleotides along with the number of references in the standard VCF format (pg. 2, col. 2, para. 2, lines 4-11; pg. 9, col. 2, para. 1, lines 2-4; pg. 10-12, Table 2). This teaches a number of citation times. It would be obvious to calculate a citation factor based on citation times because Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Therefore, a citation factor based on citation times provides reliability in variant interpretation.
It would be obvious to calculate a risk score based on the genetic factor disclosed by Polfus et al. and Folkersen et al. and the citation factor disclosed by Chunn et al. because the results of the study disclosed by Polfus et al. are important for risk characterization in diverse populations and help to improve genetic-based risk prediction across populations (pg. 1, Summary, lines 10-12). Therefore, incorporating the elements discussed by Polfus et al. in the genetic factor helps improve genetic-based risk prediction. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the genetic factor helps advance public perception of genomic risk. Chunn et al. discloses that having ready access to the most complete database of published variants and all the associated evidence annotations is essential in reducing the time it takes to interpret a variant and ensuring the accuracy of that interpretation (pg. 2, col. 1, para. 1, lines 13-20). Also, further discloses that citation counts is a measure of relevance (pg. 5, col. 2, para. 1, lines 1-6). Therefore, a citation factor provides reliability in variant interpretation and can help determine impact of citations on variants associated with disease risk.
With respect to claim 8:
Chunn et al., Folkersen et al., and Sadat et al. do not disclose wherein for each of the M number of original parameter sets, the statistics include: a p-value representing a probability that an association of the specific disease with the corresponding one of the M number of specific risk alleles is due to random chance.
However, Polfus et al. discloses observing statistically significant associations in 39 of the 582 known risk loci in the multiethnic meta-analysis, which includes rs35011184 near TCF7L2 (MIM: 602228) (OR = 1.31, 95% confidence interval [CI] = 1.27–1.34, p = 3.32
×
10
-
102
), which was also the most significant variant in European (OR = 1.34, 95% CI = 1.29–1.38, p = 1.55
×
10
-
75
), and African populations (OR = 1.29, 95% CI = 1.20–1.39, p = 2.42
×
10
-
11
) (pg. 4, col. 2, para. 1, lines 7-25). Also, further discloses that summary statistics from each contributing GWAS were combined across studies to form a single combined log odds ratio (OR) estimate, standard error, and Wald test for each variant (pg. 3, col. 2, para. 4, lines 3-7). This teaches p-values as probabilities that risk loci are associated with type 2 diabetes. The lower the p-value, the more statistically significant the associations are.
Chunn et al., Folkersen et al., and Sadat et al. do not disclose an odds ratio, which is a ratio of a probability of a person with the corresponding one of the M number of specific risk alleles getting the specific disease to a probability of a person without the corresponding one of the M number of specific risk alleles getting the specific disease.
However, Polfus et al. discloses odds ratios calculated for each risk/other allele in Table 1 (pg. 5, Table 1). Also, further discloses that summary statistics from each contributing GWAS were combined across studies to form a single combined log odds ratio (OR) estimate, standard error, and Wald test for each variant (pg. 3, col. 2, para. 4, lines 3-7). This teaches odds ratios as a ratio of a person with risk alleles getting type 2 diabetes to the person without the risk allele getting type 2 diabetes.
It would have been prima facie obvious to one of ordinary skill in the art to modify the multiethnic GWAS study disclosed by Polfus et al. to incorporate citation counts disclosed by Chunn et al., number of chromosomes disclosed by Folkersen et al., and the USB connector disclosed by Sadat et al. One would be motivated to incorporate citation counts, chromosome counts, and the USB connector because Chunn et al. discloses that Mastermind demonstrated a sensitivity of 98.4% compared to 4.4, 45.6, and 37.4% for alternatives PubMed, Google Scholar, and ClinVar (pg. 1, Abstract, lines 15-17). Therefore, incorporating citation counts from Mastermind in the multiethnic GWAS study is very reliable. Folkersen et al. discloses that the Impute.me tool utilizing the top-SNP calculation mode for polygenic score calculation focuses on education and information on common, complex disorders with polygenetic architecture, allowing the user to consider genetic risk scores in a medical context relevant to the individual (pg. 1-2, Abstract, lines 19-26). Therefore, incorporating the number of chromosomes in the multiethnic GWAS study helps advance public perception of genomic risk. Sadat et al. discloses that Type-C USB connectors deliver more power than any previous USB version (pg. 4, col. 1, para. 1, lines 1-2). Incorporating the USB connector in the multiethnic GWAS study will improve efficiency and power. There is a likelihood of success, since all teachings are of either genetic variant analyses and disease risk scoring, or USB Type-C, which are well known elements in the field of bioinformatics.
Claims Free from Prior Art
Claims 3 and 9 recite wherein the step of calculating a genetic factor is to calculate the genetic factor according to a formula:
F
a
c
t
o
r
G
e
n
e
t
i
c
=
1
M
∑
i
=
1
N
-
log
10
P
i
×
O
R
i
×
S
N
P
_
T
y
p
e
i
×
F
a
c
t
o
r
R
a
c
e
,
i
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
e
f
,
i
, where
F
a
c
t
o
r
G
e
n
e
t
i
c
represents the genetic factor,
P
i
represents the p-value for an
i
t
h
one of the N number of specific risk alleles,
O
R
i
represents the odds ratio for the
i
t
h
one of the N number of specific risk alleles,
S
N
P
_
T
y
p
e
i
represents the number of chromosomes in a homologous chromosome pair having the
i
t
h
one of the N number of specific risk alleles,
F
a
c
t
o
r
R
a
c
e
,
i
represents the race factor for the
i
t
h
one of the N number of specific risk alleles, and
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
e
f
,
i
represents the global reference allele frequency for the
i
t
h
one of the N number of specific risk alleles. Claims 4 and 10 recite wherein for an
i
t
h
one of the target parameter sets that corresponds to an
i
t
h
one of the N number of specific risk alleles, i being an integer ranging from one to N, the step of calculating a race factor is to calculate the race factor according to formulas:
F
a
c
t
o
r
R
a
c
e
,
i
=
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
+
1
,
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
≥
0
1
1
-
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
,
log
10
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
<
0
; and
F
r
e
q
u
e
n
c
y
_
r
a
t
i
o
G
r
o
u
p
r
i
s
k
,
i
=
F
r
e
q
u
e
n
c
y
G
r
o
u
p
r
i
s
k
,
i
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
i
s
k
,
i
, where
F
a
c
t
o
r
R
a
c
e
,
i
represents the race factor for the
i
t
h
one of the N number of specific risk alleles,
F
r
e
q
u
e
n
c
y
G
r
o
u
p
r
i
s
k
,
i
represents the group-specific risk allele frequency for the
i
t
h
one of the N number of specific risk alleles, and
F
r
e
q
u
e
n
c
y
G
l
o
b
a
l
r
i
s
k
,
i
represents the global risk allele frequency for the
i
t
h
one of the N number of specific risk alleles. Claims 5 and 11 recite wherein the step of calculating a citation factor is to calculate the citation factor according to a formula:
F
a
c
t
o
r
c
i
t
a
t
i
o
n
=
ln
∑
i
=
1
N
(
C
i
t
a
t
i
o
n
n
u
m
i
+
1
)
, where
F
a
c
t
o
r
c
i
t
a
t
i
o
n
represents the citation factor, and
C
i
t
a
t
i
o
n
n
u
m
i
represents the number of citation times for an
i
t
h
one of the N number of specific risk alleles. Claims 6 and 12 recite wherein the step of calculating a risk score is to calculate the risk score according to a formula:
S
c
o
r
e
r
i
s
k
=
100
,
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
>
100
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
,
0
<
F
a
c
t
o
r
G
e
n
e
t
i
c
×
F
a
c
t
o
r
C
i
t
a
t
i
o
n
≤
100
, where
S
c
o
r
e
r
i
s
k
represents the risk score,
F
a
c
t
o
r
G
e
n
e
t
i
c
represents the genetic factor, and
F
a
c
t
o
r
C
i
t
a
t
i
o
n
represents the citation factor. These limitations are free of the art.
Conclusion
No claims are allowed.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jammy Luo whose telephone number is (571)272-2358. The examiner can normally be reached Monday - Friday, 9:00 AM - 5:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry D Riggs can be reached at (571)270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.N.L./Examiner, Art Unit 1686
/OLIVIA M. WISE/Supervisory Patent Examiner, Art Unit 1685