US2015205914A1PendingUtilityA1

System and Methods for Detecting Genetic Variation

Assignee: COUNSYL INCPriority: Oct 31, 2012Filed: Oct 10, 2014Published: Jul 23, 2015
Est. expiryOct 31, 2032(~6.3 yrs left)· nominal 20-yr term from priority
C12Q 1/6874G06F 19/18G06F 19/22G16B 20/20G16B 20/40G16B 30/10G16B 20/10G16B 30/00G16B 20/00C40B 30/00
63
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The invention provides methods, apparatuses, and compositions for high-throughput amplification sequencing of specific target sequences in one or more samples. In some aspects, barcode-tagged polynucleotides are sequenced simultaneously and sample sources are identified on the basis of barcode sequences. In some aspects, sequencing data are used to determine one or more genotypes at one or more loci comprising a causal genetic variant. In some aspects, systems and methods of detecting genetic variation are provided.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 - 60 . (canceled) 
     
     
         61 . A method of detecting genetic variation in a subject's genome comprising:
 (a) providing a plurality of clusters of polynucleotides, wherein (i) each cluster comprises multiple copies of a nucleic acid duplex attached to a support; (ii) each duplex in a cluster comprises a first molecule comprising sequences A-B-G′-D′-C′ from 5′ to 3′ and a second molecule comprising sequences C-D-G-B′-A′ from 5′ to 3′; (iii) sequence A′ is complementary to sequence A, sequence B′ is complementary to sequence B, sequence C′ is complementary to sequence C, sequence D′ is complementary to sequence D, and sequence G′ is complementary to sequence G; (iv) sequence G is a portion of a target polynucleotide sequence from a subject and is different for each of a plurality of clusters; (v) sequence B′ is located 5′ with respect to sequence G in the corresponding target polynucleotide sequence; and (vi) each first molecule comprises a barcode sequence;   (b) sequencing sequence G′ by extension of a first primer comprising sequence D to produce an R1 sequence for each cluster;   (c) sequencing sequence B′ by extension of a second primer comprising sequence A to produce R2 sequence for each cluster;   (d) hybridizing a third primer to sequence C′ and sequencing the barcode sequence by extension of the third primer to produce a barcode sequence for each cluster;   (e) performing a first alignment using a first algorithm to align all R1 sequences to a first reference sequence;   (f) performing a second alignment using a second algorithm to locally align R1 sequences identified in said first alignment as likely to contain an insertion or deletion with respect to the first reference sequence, to produce a single consensus alignment for each insertion or deletion;   (g) performing an R2 alignment by aligning all R2 sequences to a second reference sequence; and   (h) determining the presence or absence of a sequence variation identified by steps (e) to (g).   
     
     
         62 . The method of  claim 61 , further comprising transmitting a report with the sequence variation determined in step (h) to a receiver. 
     
     
         63 . The method of  claim 61 , wherein the first reference sequence comprises a reference genome. 
     
     
         64 . The method of  claim 61 , wherein the second reference sequence consists of every sequence B for every different target polynucleotide. 
     
     
         65 . The method of  claim 61 , wherein R2 sequences are aligned independently of R1 sequences. 
     
     
         66 . The method of  claim 61 , further comprising discarding an R1 sequence that aligns to a first position in the first reference sequence that is more than 10,000 base pairs away from a second position in the first reference sequence to which the R2 sequence for the same cluster aligns. 
     
     
         67 . The method of  claim 61 , further comprising deleting a portion of an R1 sequence for a cluster when the portion of R1 sequence to be deleted is identical to at least a portion of sequence B′ for that cluster and sequence G is shorter than the R1 sequence for that cluster. 
     
     
         68 . The method of  claim 61 , further comprising deleting a portion of an R1 sequence for a cluster when the portion of R1 sequence to be deleted is identical to at least a portion of any sequence B′, the portion includes either the 5′ or 3′ nucleotide of R1, and either (i) no R2 sequence was produced for the cluster or (ii) R2 sequence produced is not identical to any sequence B. 
     
     
         69 . The method of  claim 61 , wherein performing the first alignment with a system using the first algorithm takes less time to align all R1 reads than would be taken if the system used the second algorithm to perform the first alignment. 
     
     
         70 . The method of  claim 61 , wherein performing the first alignment with a system using the first algorithm uses less system memory to align all R1 reads than would be used if the system used the second algorithm to perform the first alignment. 
     
     
         71 . The method of  claim 61 , wherein said first algorithm is based on Burrows-Wheeler transform. 
     
     
         72 . The method of  claim 61 , wherein said second algorithm is based on Smith-Waterman algorithm or a hash function. 
     
     
         73 . The method of  claim 61 , wherein R1 and R2 sequences are generated for at least 100 different target polynucleotides. 
     
     
         74 . The method of  claim 61 , wherein each barcode differs from every other barcode in a plurality of different barcodes analyzed in parallel. reaction. 
     
     
         75 . The method of  claim 61 , wherein the barcode sequence is located 5′ from sequence D′. 
     
     
         76 . The method of  claim 61 , further comprising grouping sequences from the clusters based on the barcode sequences. 
     
     
         77 . The method of  claim 75 , further comprising discarding all but one of a plurality of R1 sequences having the same sequence and alignment within a barcode sequence grouping. 
     
     
         78 . The method of  claim 61 , wherein sequences A, B, C, and D are at least 5 nucleotides in length. 
     
     
         79 . The method of  claim 61 , wherein sequence G of every cluster is 1 to 1000 nucleotides in length. 
     
     
         80 . The method of  claim 61 , wherein each probe sequence B of a plurality of clusters is complementary to a sequence comprising a causal genetic variant or a sequence within 200 nucleotides of a causal genetic variant. 
     
     
         81 . The method of  claim 61 , wherein an R1 sequence is produced for at least about 10 8  clusters in a single reaction. 
     
     
         82 . The method of  claim 61 , wherein presence, absence, or allele ratio of one or more causal genetic variants is determined with an accuracy of at least about 90%. 
     
     
         83 . The method of  claim 61 , wherein the consensus sequence identifies an insertion, a deletion, or an insertion and a deletion in a target polynucleotide with an accuracy of at least about 90%. 
     
     
         84 . The method of  claim 61 , wherein each probe sequence B of a plurality of clusters is complementary to a sequence comprising a non-subject sequence or a sequence within 200 nucleotides of a non-subject sequence. 
     
     
         85 . The method of  claim 61 , wherein the presence or absence of one or more non-subject sequences is determined with an accuracy of at least about 90%. 
     
     
         86 . The method of  claim 61 , further comprising calculating a plurality of probabilities based on the R1 sequences for the subject and including the probabilities in the report, wherein each probability is a probability of the subject or a subject's offspring having or developing a disease or trait. 
     
     
         87 . A method of detecting genetic variation in a subject's genome comprising:
 (a) providing a plurality of clusters of polynucleotides, wherein (i) each cluster comprises multiple copies of a nucleic acid duplex attached to a support; (ii) each duplex in a cluster comprises a first molecule comprising sequences A-B-G′-D′-C′ from 5′ to 3′ and a second molecule comprising sequences C-D-G-B′-A′ from 5′ to 3′, (iii) sequence A′ is complementary to sequence A, sequence B′ is complementary to sequence B, sequence C′ is complementary to sequence C, sequence D′ is complementary to sequence D, and sequence G′ is complementary to sequence G; (iv) sequence G is a portion of a target polynucleotide sequence from a subject and is different for each of a plurality of clusters; (v) sequence B′ is located 5′ with respect to sequence G in the corresponding target polynucleotide sequence; and (vi) each first molecule comprises a barcode sequence;   (b) sequencing sequence G′ by extension of a first primer comprising sequence D to produce an R1 sequence for each cluster;   (c) sequencing sequence B′ by extension of a second primer comprising sequence A to produce R2 sequence for each cluster;   (d) further comprising hybridizing a third primer to sequence C′ and sequencing the barcode sequence by extension of the third primer to produce a barcode sequence for each cluster   (e) performing a first alignment using a first algorithm to align all R1 sequences to a first reference sequence;   (f) performing a second alignment using a second algorithm to locally align R1 sequences identified in said first alignment as likely to contain an insertion or deletion with respect to the first reference sequence, to produce a single consensus alignment for each insertion or deletion;   (g) performing an R2 alignment by aligning all R2 sequences to a second reference sequence; and   (h) calculating a plurality of probabilities based on the R1 sequences for the subject and including the probabilities in a report identifying sequence variation identified by steps (e) to (g), wherein each probability is a probability of the subject or a subject's offspring having or developing a disease or trait.   
     
     
         88 . The method of  claim 87 , further comprising transmitting a report with the sequence variation determined in step (h) to a receiver. 
     
     
         89 . The method of  claim 87 , wherein the first reference sequence comprises a reference genome. 
     
     
         90 . The method of  claim 87 , wherein the second reference sequence consists of every sequence B for every different target polynucleotide. 
     
     
         91 . The method of  claim 87 , wherein R2 sequences are aligned independently of R1 sequences. 
     
     
         92 . The method of  claim 87 , further comprising discarding an R1 sequence that aligns to a first position in the first reference sequence that is more than 10,000 base pairs away from a second position in the first reference sequence to which the R2 sequence for the same cluster aligns. 
     
     
         93 . The method of  claim 87 , further comprising deleting a portion of an R1 sequence for a cluster when the portion of R1 sequence to be deleted is identical to at least a portion of sequence B′ for that cluster and sequence G is shorter than the R1 sequence for that cluster. 
     
     
         94 . The method of  claim 87 , further comprising deleting a portion of an R1 sequence for a cluster when the portion of R1 sequence to be deleted is identical to at least a portion of any sequence B′, the portion includes either the 5′ or 3′ nucleotide of R1, and either (i) no R2 sequence was produced for the cluster or (ii) R2 sequence produced is not identical to any sequence B. 
     
     
         95 . The method of  claim 87 , wherein performing the first alignment with a system using the first algorithm takes less time to align all R1 reads than would be taken if the system used the second algorithm to perform the first alignment. 
     
     
         96 . The method of  claim 87 , wherein performing the first alignment with a system using the first algorithm uses less system memory to align all R1 reads than would be used if the system used the second algorithm to perform the first alignment. 
     
     
         97 . The method of  claim 87 , wherein said first algorithm is based on Burrows-Wheeler transform. 
     
     
         98 . The method of  claim 87 , wherein said second algorithm is based on Smith-Waterman algorithm or a hash function. 
     
     
         99 . The method of  claim 87 , wherein R1 and R2 sequences are generated for at least 100 different target polynucleotides. 
     
     
         100 . The method of  claim 87 , wherein each barcode differs from every other barcode in a plurality of different barcodes analyzed in parallel. 
     
     
         101 . The method of  claim 87 , wherein the barcode sequence is associated with a single sample in a pool of samples sequenced in a single reaction. 
     
     
         102 . The method of  claim 87 , wherein each of a plurality of barcode sequences is uniquely associated with a single sample in a pool of samples sequenced in a single reaction. 
     
     
         103 . The method of  claim 87 , wherein the barcode sequence is located 5′ from sequence D′. 
     
     
         104 . A method of detecting genetic variation in a subject's genome comprising:
 (a) providing a plurality of clusters of polynucleotides, wherein (i) each cluster comprises multiple copies of a nucleic acid duplex attached to a support; (ii) each duplex in a cluster comprises a first molecule comprising sequences A-B-G′-D′-C′ from 5′ to 3′ and a second molecule comprising sequences C-D-G-B′-A′ from 5′ to 3; (iii) sequence A′ is complementary to sequence A, sequence B′ is complementary to sequence B, sequence C′ is complementary to sequence C, sequence D′ is complementary to sequence D, and sequence G′ is complementary to sequence G; (iv) sequence G is a portion of a target polynucleotide sequence from a subject and is different for each of a plurality of clusters; (v) sequence B′ is located 5′ with respect to sequence G in the corresponding target polynucleotide sequence; and (vi) each first molecule comprises a barcode sequence;   (b) sequencing sequence G′ by extension of a first primer comprising sequence D to produce an R1 sequence for each cluster;   (c) sequencing sequence B′ by extension of a second primer comprising sequence A to produce R2 sequence for each cluster;   (d) hybridizing a third primer to sequence C′ and sequencing the barcode sequence by extension of the third primer to produce a barcode sequence for each cluster;   (e) further comprising grouping sequences from the clusters based on the barcode sequences;   (f) performing a first alignment using a first algorithm to align all R1 sequences to a first reference sequence;   (g) performing a second alignment using a second algorithm to locally align R1 sequences identified in said first alignment as likely to contain an insertion or deletion with respect to the first reference sequence, to produce a single consensus alignment for each insertion or deletion;   (h) performing an R2 alignment by aligning all R2 sequences to a second reference sequence; and   (i) calculating a plurality of probabilities based on the R1 sequences for the subject and including the probabilities in a report identifying sequence variation identified by steps (f) to (h), wherein each probability is a probability of the subject or a subject's offspring having or developing a disease or trait.   
     
     
         105 . A method of detecting genetic variation in a subject's genome comprising:
 (a) providing a plurality of clusters of polynucleotides, wherein (i) each cluster comprises multiple copies of a nucleic acid duplex attached to a support; (ii) each duplex in a cluster comprises a first molecule comprising sequences A-B-G′-D′-C′ from 5′ to 3′ and a second molecule comprising sequences C-D-G-B′-A′ from 5′ to 3′; (iii) sequence A′ is complementary to sequence A, sequence B′ is complementary to sequence B, sequence C′ is complementary to sequence C, sequence D′ is complementary to sequence D, and sequence G′ is complementary to sequence G; (iv) sequence G is a portion of a target polynucleotide sequence from a subject and is different for each of a plurality of clusters; (v) sequence B′ is located 5′ with respect to sequence G in the corresponding target polynucleotide sequence; and (vi) each first molecule comprises a barcode sequence;   (b) sequencing sequence G′ by extension of a first primer comprising sequence D to produce an R1 sequence for each cluster;   (c) sequencing sequence B′ by extension of a second primer comprising sequence A to produce R2 sequence for each cluster;   (d) hybridizing a third primer to sequence C′ and sequencing the barcode sequence by extension of the third primer to produce a barcode sequence for each cluster   (e) grouping sequences from the clusters based on the barcode sequences; and   (f) further comprising discarding all but one of a plurality of R1 sequences having the same sequence and alignment within a barcode sequence grouping;   (g) performing a first alignment using a first algorithm to align all R1 sequences to a first reference sequence;   (h) performing a second alignment using a second algorithm to locally align R1 sequences identified in said first alignment as likely to contain an insertion or deletion with respect to the first reference sequence, to produce a single consensus alignment for each insertion or deletion;   (i) performing an R2 alignment by aligning all R2 sequences to a second reference sequence; and   (j) calculating a plurality of probabilities based on the R1 sequences for the subject and including the probabilities in a report identifying sequence variation identified by steps (d) to (f), wherein each probability is a probability of the subject or a subject's offspring having or developing a disease or trait.   
     
     
         106 . The method of  claim 87 , wherein sequences A, B, C, and D are at least 5 nucleotides in length. 
     
     
         107 . The method of  claim 87 , wherein sequence G of every cluster is 1 to 1000 nucleotides in length. 
     
     
         108 . The method of  claim 87 , wherein each probe sequence B of a plurality of clusters is complementary to a sequence comprising a causal genetic variant or a sequence within 200 nucleotides of a causal genetic variant. 
     
     
         109 . The method of  claim 87 , wherein sequence B of one or more of the clusters comprises a sequence selected from the group consisting of SEQ ID NOs:22-121. 
     
     
         110 . The method of  claim 87 , wherein an R1 sequence is produced for at least about 10 8  clusters in a single reaction. 
     
     
         111 . The method of  claim 87 , wherein presence, absence, or allele ratio of one or more causal genetic variants is determined with an accuracy of at least about 90%. 
     
     
         112 . The method of  claim 87 , wherein the consensus sequence identifies an insertion, a deletion, or an insertion and a deletion in a target polynucleotide with an accuracy of at least about 90%. 
     
     
         113 . The method of  claim 87 , wherein each probe sequence B of a plurality of clusters is complementary to a sequence comprising a nonsubject sequence or a sequence within 200 nucleotides of a non-subject sequence. 
     
     
         114 . The method of  claim 87 , wherein the presence or absence of one or more non-subject sequences is determined with an accuracy of at least about 90%. 
     
     
         115 . The method of  claim 104  further comprising transmitting the report of step (i) to a receiver. 
     
     
         116 . The method of  claim 105  further comprising transmitting the report of step (j) to a receiver. 
     
     
         117 . A method of detecting genetic variation in a subject's genome comprising:
 a. Providing a plurality of fragmented polynucleotides from the subject;   b. Ligating a first partially single-stranded adapter to the polynucleotides of step (a), wherein the partially single-stranded adapter has a double-stranded region at one end (sequence U hybridized to complementary sequence U′) and the single-stranded sequence Y that does not hybridize to the target polynucleotide under the hybridization and extension conditions used, and further wherein ligation adds sequence Y to both 5′ ends of the target polynucleotides;   c. Hybridizing a plurality of a plurality of different oligonucleotide primers, each having a different target-specific sequence W at the 3′ end and extending the primers to produce an extended oligonucleotide with sequence Y′ (complement of Y) at the 3′ end;   d. Optionally, amplifying the extended oligonucleotides with a pair of amplification primers comprising:
 i. A first amplification primer comprising sequence X and sequence Y, with sequence Y at the 3′ end for hybridization to sequence Y′; 
 ii. A second amplification primer comprising sequences V and Z, with Z at the 3′ end for hybridization to sequence Z′ of an extended X-Y primer; 
    Wherein amplification produces a plurality of extended X-Y oligonucleotides comprising sequences X, Y, W′, and Z′ (5′ to 3′; where W′ is the complement of W, and Z′ is the complement of Z) from the first amplification primer, and a plurality of sequences comprising V, Z, Y′, and X′ (5′ to 3′; where X′ is the complement of X) from the second amplification primer;   e. Sequencing a plurality of different target polynucleotides, each contained in a polynucleotide comprising one strand comprising sequences V, Z, W, Y′, and X′ (from 5′ to 3′), and another strand comprising sequences X, Y, W′, Z′, and V′ (from 5′ to 3′), with target polynucleotide sequence located between Z/Y′ and between Z′/Y;   f. Determining if the subject has a genetic variation based on the sequencing of step (e).   
     
     
         118 . The method of  claim 117 , wherein each oligonucleotide primer of step (c) further comprises a binding partner at the 5′ end. 
     
     
         119 . The method of  claim 117 , wherein the extended oligonucleotides are amplified. 
     
     
         120 . The method of  claim 119 , wherein the extended oligonucleotides are exponentially amplified. 
     
     
         121 . The method of  claim 117 , wherein sequence Z is common among all oligonucleotide primers. 
     
     
         122 . The method of  claim 117 , wherein sequence W is different for each different oligonucleotide primer, is positioned at the 3′ end of each oligonucleotide primer, and is complementary to a sequence comprising a causal genetic variant or a sequence within 200 nucleotides of a causal genetic variant. 
     
     
         123 . The method of  claim 117 , wherein one or more of sequences V, W, X, Y, and Z are different sequences. 
     
     
         124 . The method of  claim 117 , wherein one or more of sequences V, W, X, Y, and Z comprise 5 or more nucleotides each. 
     
     
         125 . The method of  claim 117 , wherein sequence W of one or more of the plurality of oligonucleotide primers comprises a sequence selected from the group consisting of SEQ ID NOs 22-121. 
     
     
         126 . The method of  claim 117 , wherein the fragmented polynucleotides have a median length between about 200 and about 1000 base pairs. 
     
     
         127 . The method of  claim 117 , wherein sequence Y is positioned at the 3′ end of the adapter. 
     
     
         128 . The method of  claim 117 , wherein the fragmented polynucleotides are treated to produce blunt ends or to have a defined overhang prior to step (a). 
     
     
         129 . The method of  claim 128 , wherein the overhang consists of an adenine. 
     
     
         130 . The method of  claim 117 , further comprising transmitting a report identifying the sequence variation determined in step (f) to a receiver.

Join the waitlist — get patent alerts

Track US2015205914A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.