Genetic diagnosis using multiple sequence variant analysis
Abstract
The present invention is in the field of nucleic acid-based genetic analysis. More particularly, it discloses novel insights into the overall structure of genetic variation in all living species. The structure can be revealed with the use of any data set of genetic variants from a particular locus. The invention is useful to define the subset of variations that are most suited as genetic markers to search for correlations with certain phenotypic traits. Additionally, the insights are useful for the development of algorithms and computer programs that convert genotype data into the constituent haplotypes that are laborious and costly to derive in an experimental way. The invention is useful in areas such as (i) genome-wide association studies, (ii) clinical in vitro diagnosis, (iii) plant and animal breeding, (iv) the identification of micro-organisms.
Claims
exact text as granted — not AI-modified1 - 2 . (canceled)
3 . A method of preparing an SPC network of a genomic region of interest comprising the steps of:
a. obtaining the nucleic acid sequence of said genomic region of interest from a plurality of subjects; b. identifying a plurality of polymorphisms in said nucleic acid sequences; c. identifying the haploid genotypes of said polymorphisms in said nucleic acid sequences; d. computing the pairwise haploid genotypes in said nucleic acid sequences for each combination of two polymorphisms of said polymorphisms by combining for each of said nucleic acid sequences the genotype of the first polymorphism with the genotype of the second polymorphism; e. assigning a polymorphism as belonging to an SPC network if the pairwise haploid genotypes for each combination of the polymorphism with each of the polymorphisms of the SPC network comply with one of the following rules:
(i) two polymorphisms are defined as belonging to one SPC in the network when they exhibit only the major (AA) and minor (BB) pairwise haploid genotypes;
(ii) two polymorphisms are defined as having a dependent relationship with each other in the network when they exhibit only the major (AA), minor (BB) and mixed (BA) pairwise haploid genotypes or when they exhibit only the major (AA), minor (BB) and mixed (AB) pairwise haploid genotypes;
(iii) two polymorphisms are defined as having an independent relationship with each other in the network when they exhibit only the major (AA), mixed (AB) and mixed (BA) pairwise haploid genotypes,
f. compiling an SPC network by repeating step (e) until the largest possible number of polymorphisms from the genomic region of interest have been assigned to said SPC network.
4 . A method of preparing an SPC network of a genomic region of interest comprising the steps of:
a. obtaining the nucleic acid sequence of said genomic region of interest from a plurality of subjects; b. identifying a plurality of polymorphisms in said nucleic acid sequences; c. identifying the diploid genotypes of said polymorphisms in said nucleic acid sequences; d. computing the pairwise diploid genotypes in said nucleic acid sequences for each combination of two polymorphisms of said polymorphisms by combining for each of said nucleic acid sequences the genotype of the first polymorphism with the genotype of the second polymorphism; e. assigning a polymorphism as belonging to an SPC network if the pairwise diploid genotypes for each combination of the polymorphism with each of the polymorphisms of the SPC network comply with one or more of the following rules:
(i) two polymorphisms are defined as belonging to one SPC in the network when they exhibit only the homozygous major (AA), homozygous minor (BB) and heterozygous (HH) pairwise diploid genotypes;
(ii) two polymorphisms are defined as having a dependent relationship with each other in the network when they exhibit only the homozygous major (AA), mixed (AH), mixed (AB), heterozygous (HH), mixed (HB) and homozygous minor (BB) pairwise diploid genotypes, or they exhibit only the homozygous major (AA), mixed (HA), mixed (BA), heterozygous (HH), mixed (BH), and homozygous minor (BB) pairwise diploid genotypes;
(iii) two polymorphisms are defined as having a independent relationship with each other in the network when they exhibit only the homozygous major (AA), mixed (AH), mixed (HA), mixed (AB), mixed (BA) and heterozygous (HH) pairwise diploid genotypes,
f. compiling an SPC network by repeating step (e) until the largest possible number of polymorphisms from the genomic region of interest have been assigned to said SPC network.
5 . A method of determining the SPC-haplotypes from unphased diploid genotypes of a genomic region of interest of a subject, comprising:
a. obtaining an SPC network according to claim 4 ; b. determining the SPC-haplotypes from said SPC network according to the following rules:
(i) each SPC defines a separate SPC-haplotype unless the minor allele of the polymorphism(s) of said SPC always coincide with the minor allele of the polymorphism(s) of the SPCs that depend from said SPC, and wherein the separate SPC-haplotype is defined by the minor alleles of the polymorphisms of said SPC, the minor alleles of all the polymorphisms of the SPCs from which said SPC depends, and the major alleles of all remaining polymorphisms; and
(ii) a haplotype that comprises the major allele at all polymorphic sites that are part of the SPC network, is present in case the total sum of the number of occurrences of all the independent SPCs is lower than the total number of haplotypes in the plurality of subjects;
c. identifying which SPC-haplotype or combination of two SPC-haplotypes accounts for the observed diploid genotype of said subject.
6 . A method of identifying genotyping errors in diploid SNP genotypes from a number of different subjects, comprising:
a. obtaining an SPC network according to claim 4 ; b. identifying a pair of overlapping SNPs and/or SPCs wherein the SNP or at least one SNP of one SPC falls within the boundaries of the other SPC in the pair; c. identifying pairs of said overlapping SNPs and/or SPCs which do not comply with the rules (ii) and (iii) of claim 4 , thereby detecting SNPs and/or SPCs having one or more genotyping errors; d. identifying the pairwise diploid genotypes that comprise genotyping errors in said SNPs and/or SPCs by counting the pairwise diploid genotypes in each of the following three sets:
(i) a set of pairwise diploid genotypes that is comprised of the mixed (AH), the mixed (AB) and the mixed (HB) pairwise diploid genotypes, or
(ii) a set of pairwise diploid genotypes that is comprised of the mixed (HA), the mixed (BA) and the mixed (BH) pairwise diploid genotypes, or
(iii) a set of pairwise diploid genotypes that is comprised of the mixed (BH), the mixed (HB) and the homozygous minor (BB) pairwise diploid genotypes, and
identifying which of the sets (d)(i), (d)(ii) or (d)(iii) has the lowest number of pairwise diploid genotypes to identify the set of pairwise diploid genotypes that have errors in said SNPs and/or SPCs e. identifying the subjects that have the set of diploid genotypes identified in step (d) as having errors in their SNPs and/or SPCs f. determining whether the error observed in said subjects resides in the SNP or the SPC by assigning the genotyping errors in said subjects according to the following rules:
(i) in the case of overlapping SPCs the genotyping errors are assigned to the SPC which has the fewest number of SNPs
(ii) in the case of overlapping SNPs and SPCs the genotyping errors are assigned to the SNP, unless the genotypes in the subject(s) of the SPC were already identified as errors under (f)(i).
7 - 10 . (canceled)
11 . A method of producing an SPC map of a genomic region of interest comprising the steps of:
a. obtaining the nucleic acid sequence of said genomic region of interest from a plurality of subjects; b. identifying a plurality of polymorphisms in said nucleic acid sequences; c. identifying one or more SPCs, wherein each SPC comprises a subset of polymorphisms from said nucleic acid sequence wherein said polymorphisms of said subset coincide with each other polymorphism of said subset; and d. identifying polymporphisms that do not coincide with any other polymorphism but do cosegregate with at least one SPC.
12 . A method of producing an SPC map of a genomic region of interest from unphased diploid genotypes comprising the steps of:
a. obtaining the unphased diploid genotypes of a genomic region of interest from a plurality of subjects; b. determining the major and minor metatypes found in said unphased diploid genotypes; c. identifying one or more SPCs, wherein each SPC comprises a subset of polymorphisms from said metatypes wherein said polymorphisms of said subset coincide with each other polymorphism of said subset; and d. identifying polymporphisms that do not coincide with any other polymorphism but do cosegregate with at least one SPC.
13 . A method of producing an SPC map of a genomic region of interest from the genotypes of sample pools comprising the steps of:
a. obtaining the genotypes of a genomic region of interest from a plurality of sample pools; b. determining the major and minor metatypes found in said genotypes; c. identifying one or more SPCs, wherein each SPC comprises a subset of polymorphisms from said metatypes wherein said polymorphisms of said subset coincide with each other polymorphism of said subset.
14 . The method of claim 11 or 12 , wherein said identifying one or more SPCs comprises identifying each polymorphism of said subset that coincides with each other polymorphism of said subset according to a percentage coincidence of the minor alleles of said polymorphisms of between 75% and 100%.
15 . The method of claim 11 or 12 , wherein said identifying one or more SPCs comprises multiple rounds of coincidence analysis.
16 . The method of claim 11 or 12 , wherein each successive round of coincidence analysis is performed at a decreasing percentage coincidence from 100% coincidence to 75% coincidence.
17 . The method of claim 11 or 12 , wherein the coincidence of each said polymorphism of said subset with each other polymorphism of said subset is calculated according to a parameter selected from the group consisting of a pairwise C value, C* value, a r 2 linkage disequilibrium value, a Δ linkage disequilibrium value, a δ linkage disequilibrium value, and a d linkage disequilibrium value.
18 . The method of claim 17 , wherein said parameter is a pairwise C value of from 0.75 to 1.
19 . The method of claim 11 or 12 , wherein the identification of a plurality of polymorphisms in said target nucleic acid sequences is determined by an assay selected from the group consisting of direct sequence analysis, differential nucleic acid analysis, sequence based genotyping, DNA chip analysis, and polymerase chain reaction analysis.
20 . A method of selecting one or more polymorphisms from a genomic region of interest for use in genotyping, comprising the steps of:
a. obtaining an SPC map according to claim 11 or 12 ; b. selecting at least one cluster tag polymorphism which identifies a specific SPC in said SPC map; and c. selecting a sufficient number of cluster tag polymorphisms for use in a genotyping study of the genomic region of interest.
21 . The method of claim 20 , wherein said cluster tag polymorphism is selected from the group consisting of a single nucleotide polymorphism (SNP), a deletion polymorphism, an insertion polymorphism; and a short tandem repeat polymorphism (STR).
22 . The method of claim 20 , wherein said cluster tag polymorphism is a known SNP associated with a genetic trait.
23 . A method of identifying a marker for a trait or phenotype comprising:
a. obtaining a sufficient number of cluster tag polymorphisms according to claim 20 ; b. assessing said cluster tag polymorphisms to identify an association between a trait or phenotype and at least one cluster tag polymorphism, wherein identification of said association identifies said cluster tag polymorphism as a marker for said trait or phenotype.
24 . The method of claim 23 , wherein a cluster tag polymorphism is correlated with a trait or phenotype selected from the group comprising a genetic disorder, a predisposition to a genetic disorder, susceptibility to a disease, an agronomic or livestock performance trait, a product quality trait.
25 . The method of claim 23 , wherein said marker is a marker of a genetic disorder and said SPC map is prepared according to claim 11 or claim 12 , and said plurality of subjects each manifests the same genetic disorder.
26 . The method of claim 23 , wherein the identification of a plurality of polymorphisms in said target nucleic acid sequences is determined by an assay selected from the group consisting of direct sequence analysis, differential nucleic acid analysis, sequence based genotyping, DNA chip analysis and polymerase chain reaction analysis.
27 . The method of claim 23 , comprising further identifying non-clustering polymorphisms, wherein said non-clustering polymorphisms do not co-segregate with other polymorphisms but do co-segregate with at least one SPC.
28 . A method for in vitro diagnosis of a trait or a phenotype in a subject comprising:
a. obtaining a marker for said trait or phenotype according to claim 22 ; b. obtaining a target nucleic acid sample from said subject; and c. determining the presence of said marker for said trait or a phenotype in said target nucleic acid sample, wherein the presence of said marker in said target nucleic acid indicates that said subject has the trait or the phenotype.
29 . The method of claim 28 , wherein, said trait or phenotype is selected from the group comprising a genetic disorder, a predisposition to a genetic disorder, susceptibility to a disease, an agronomic or livestock performance trait, and a product quality trait.
30 - 55 . (canceled)Join the waitlist — get patent alerts
Track US2009104601A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.