Intrinsic chromosomal linkage and disease prediction
Abstract
An improved method is the product of an exploration for ways in which to improve the success of genomic disease prediction and presents the elements of more successful prediction tools for the identification of genomic locations of disease based on: genomic labeling it loci as determined by correlated SNP linkage as determined from several additional GWAS (Genome Wide Association Studies) databases. The Wellcome Trust T2D (type 2, adult-onset, diabetes) has played a particularly important role in developing the new tools; additionally, the development of a wellness classifier, probably in part due to counter disease mutations has proven to be a powerful new tool and concept. The result has been a substantial increase in the successful genomic prediction of disease, for example, for the Wellcome Trust data there is compelling evidence of a greater than 99% successful prediction rate.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . Method of predicting complex diseases by using genomic data comprising the steps of:
a. accessing SNP data for representing a sequence of genetic material for an individual from a SNP data base; b. examining each SNP of the sequence for a pair of symbols [A, C, G. T] for associated nucleotides to unambiguously define two alleles; c. standardizing each SNP to arrange the symbols in a consistent predetermined order; d. vectorizing the symbol data to transform the SNP sequence to a digital format; e. determining the probability that a given symbol will be found in a predetermined allele; f. determining the minor symbol probability at each allele of disease and control populations; g. for each allele obtaining disease and control Shannon information I d and I c , respectively, at the same allele; h. determining the incremental information IF where IF=1 d −1 c ; i. setting a criterion level of IF>0 for high risk disease loci; j. establishing strings of genomic loci and corresponding base symbols at risk loci by selecting for maximum correlation with the disease population and minimal correlation with the control population to produce a disease classifier; k. determining the correlation coefficients of alleles on the basis of the minor symbol probabilities, for the disease and control populations; l. determining from said correlation coefficients the allele linkages, L, for the disease and control populations; m. determining a linkage classifier based on selection of alleles having both high linkage for the disease population and low linkage for the control population; n. defining a score of a sequence as being the number of coincidences of sequence symbols that the disease classifier and a sequence of symbols have in common; and o. establishing that a sequence is a disease sequence if it has a score higher than a predetermined criterion or value.
2 . A method as defined in claim 1 , wherein for the search and comparison of genomic sequences in the form of disease and control populations which finds correlated activity of individual alleles.
3 . A method as defined in claim 1 , wherein correlations are determined by the Pearson formula (for the disease matrix of rare symbols (mutants) collection, other quantities follow standard definition.
C
jk
D
=
(
M
t
D
-
μ
D
j
)
(
M
tk
D
-
μ
D
k
)
/
σ
D
j
σ
D
k
.
4 . A method as defined in claim 1 , wherein sets of alleles that exhibit correlated activity; and from this selects alleles for which disease loci are substantially greater for the disease compared to control populations.
5 . A method as defined in claim 1 , wherein incremental information is used in the selection procedure.
6 . A method as defined in claim 1 , wherein disease is predicted based on a selected set of alleles, Al, along with a set of mutant symbols W, that exhibit sufficiently greater correlated allele activity compared to the control case, and termed linkage ratio; and as well high incremental information (patent).
7 . A method as defined in claim 1 , wherein the classifier of disease is constructed in the matrix form
Classifier
=
[
Alleles
Symbols
]
,
A putative sequence then can be assigned a agreement score Sd with the classifier.
8 . A method as defined in claim 1 , wherein the disease and control populations is interchanged to select a high value loci on the basis of Incremental Information, and also leads to a wellness classifier and they wellness score, Sc.
9 . A method as defined in claim 1 , wherein the degree of disease versus wellness of a putative genomic state sequence is measured from data obtained from a patient.
10 . A method as defined in claim 10 , wherein based on assessing the disease verses wellness difference score Sd-Sc.
11 . A method as defined in claim 1 , wherein disease and wellness classifier's is obtained on the basis of the specific ancestry of the disease and control populations for a specific disease.
12 . A method as defined in claim 1 , wherein more than one ancestry disease-wellness classifier pair is assembled on a gene array for a specific disease.
13 . A method as defined in claim 1 , wherein more than one ancestry disease-wellness classifier pair of one disease is assembled on a gene array.
14 . A method as defined in claim 1 , wherein by which many ancestry disease-wellness classifiers for many diseases are assembled on a gene array.
15 . A method as defined in claim 1 , wherein the composition the probative gene arrays for a particular disease as used to assemble disease/control databases known as GWAS (Genome Wide Association Studies).
16 . A method as defined in claim 1 , wherein the study of a relatively small number of disease and control patients is correlated to the use of an extremely large number of possible alleles for the purposes of later GWAS data acquisition, whereby a far better selection procedure for choosing the alleles follows criteria based on incremental information and on disease associated snips on wellness associated sups.
17 . A method as defined in claim 1 , wherein wide ranging ancestries and diseases are assembled on a single gene array to deal for large-scale public health applications.
18 . A method as defined in claim 1 , wherein a probative gene array is determine on the basis of genome scans which searches for classifiers of high and low vulnerability to pharmaceuticals that have the risk of detrimental side effects.
19 . A method as defined in claim 1 , wherein a gene array is assembled to classify vulnerability to side effects of one pharmaceutical product.
20 . A method as defined in claim 1 , wherein the gene array is assembled to classify vulnerability to side effects of more than one pharmaceutical product.Join the waitlist — get patent alerts
Track US2018046698A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.