Methods of lowering the error rate of massively parallel dna sequencing using duplex consensus sequencing
Abstract
Next Generation DNA sequencing promises to revolutionize clinical medicine and basic research. However, while this technology has the capacity to generate hundreds of billions of nucleotides of DNA sequence in a single experiment, the error rate of approximately 1% results in hundreds of millions of sequencing mistakes. These scattered errors can be tolerated in some applications but become extremely problematic when “deep sequencing” genetically heterogeneous mixtures, such as tumors or mixed microbial populations. To overcome limitations in sequencing accuracy, a method Duplex Consensus Sequencing (DCS) is provided. This approach greatly reduces errors by independently tagging and sequencing each of the two strands of a DNA duplex. As the two strands are complementary, true mutations are found at the same position in both strands. In contrast, PCR or sequencing errors will result in errors in only one strand. This method uniquely capitalizes on the redundant information stored in double-stranded DNA, thus overcoming technical limitations of prior methods utilizing data from only one of the two strands.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A single molecule identifier adaptor molecule for use in sequencing a double-stranded target nucleic acid molecule comprising
a single molecule identifier (SMI) sequence, the SMI sequence comprising at least one degenerate or semi-degenerate nucleic acid sequence; and an SMI ligation adaptor that allows the SMI adaptor molecule to be ligated to the double-stranded target nucleic acid sequence.
2 . The single molecule identifier adaptor molecule of claim 1 , wherein the SMI sequence is single-stranded.
3 . The single molecule identifier adaptor molecule of claim 1 , wherein the SMI sequence is double-stranded.
4 . The single molecule identifier adaptor molecule of claim 1 , wherein the double-stranded target nucleic acid molecule is a double-stranded DNA or RNA molecule.
5 . The single molecule identifier adaptor molecule of claim 1 , further comprising at least two PCR primer binding sites, or at least two sequencing primer binding sites, or both.
6 . The single molecule identifier adaptor molecule of claim 1 , further comprising a double-stranded fixed reference sequence.
7 . The single molecule identifier adaptor molecule of claim 2 , wherein the degenerate or semi-degenerate nucleic acid sequence comprises a first nucleotide n-mer sequence that is between approximately 3 and 20 nucleotides in length.
8 . The single molecule identifier adaptor molecule of claim 7 , wherein the first nucleotide n-mer sequence is a degenerate sequence.
9 . The single molecule identifier adaptor molecule of claim 3 , wherein the degenerate or semi-degenerate nucleic acid sequence comprises a first nucleotide n-mer sequence, a second n-mer sequence that is complementary to the first nucleotide n-mer sequence.
10 . The single molecule identifier adaptor molecule of claim 9 , wherein the first n-mer sequence comprises a nucleotide sequence that is between approximately 3 and 20 nucleotides in length.
11 . The single molecule identifier adaptor molecule of claim 9 , wherein the first nucleotide n-mer sequence is a degenerate sequence.
12 . The single molecule identifier adaptor molecule of claim 1 , wherein the degenerate or semi-degenerate DNA sequence comprises a randomly fragmented double stranded nucleic acid derived from an alternative source.
13 . The single molecule identifier adaptor molecule of claim 1 , wherein the SMI ligation adaptor is selected from a T-overhang, an A-overhang, a CG overhang, a blunt end, or any other ligatable nucleic acid sequence.
14 . The single molecule identifier adaptor molecule of claim 1 , wherein the SMI adaptor molecule is Y-shaped, U-shaped, or a combination thereof.
15 . A method of obtaining the sequence of a double-stranded target nucleic acid comprising
ligating a double-stranded target nucleic acid molecule to at least one SMI adaptor molecule to form a double-stranded SMI-target nucleic acid complex; amplifying the double-stranded SMI-target nucleic acid complex, resulting in a set of amplified SMI-target nucleic acid products; and sequencing the amplified SMI-target nucleic acid products.
16 . The single molecule identifier adaptor molecule of claim 15 , wherein the double-stranded target nucleic acid molecule is a double-stranded DNA or RNA molecule.
17 . The method of claim 15 , further comprising generating an error-corrected double-stranded consensus sequence by (i) grouping the sequenced SMI-target nucleic acid products into families of paired target nucleic acid strands based on a common set of SMI sequences; and (ii) removing paired target nucleic acid strands having one or more nucleotide positions where the paired target nucleic acid strands disagree, or alternatively removing nucleotide positions from nucleic acid strands where the paired strands disagree at that specific position.
18 . The method of claim 15 , wherein the double-stranded target nucleic acid molecule is a sheared double-stranded DNA or RNA fragment.
19 . The method of claim 18 , wherein the sheared double-stranded nucleic acid fragment further comprises a double-stranded target nucleic acid sequence ligation adaptor.
20 . The method of claim 19 , wherein the double-stranded target nucleic acid sequence ligation adaptor is selected from a T-overhang, an A-overhang, a CG overhang, a blunt end, or any ligatable nucleic acid sequence.
21 . The method of claim 20 , wherein each end of the double-stranded target nucleic acid molecule is ligated to an SMI adaptor molecule.
22 . The method of claim 21 , wherein each SMI adaptor molecules comprises a single molecule identifier (SMI) sequence, the SMI sequence comprising a degenerate or semi-degenerate nucleic acid sequence; and
an SMI ligation adaptor that allows the SMI adaptor molecule to be ligated to the double-stranded target nucleic acid sequence.
23 . The method of claim 22 , wherein the SMI sequence is single-stranded.
24 . The method of claim 22 , wherein the SMI sequence is double-stranded.
25 . The method of claim 23 , wherein the degenerate or semi-degenerate nucleic acid sequence comprises a degenerate nucleotide n-mer sequence.
26 . The method of claim 24 , wherein the double-stranded degenerate or semi-degenerate nucleic acid sequence comprises a first nucleotide n-mer sequence and a second n-mer sequence that is complementary to the first nucleotide n-mer sequence.
27 . The method of claim 26 , wherein the first nucleotide n-mer sequence is a degenerate sequence.
28 . The method of claim 22 , further comprising at least two PCR primer binding sites, at least two sequencing primer binding sites, a double-stranded fixed reference sequence or a combination thereof.
29 . A method of generating an error corrected double-stranded consensus sequence comprising a first stage of single strand consensus sequencing (SSCS) which comprises:
tagging an individual duplex DNA molecule with an SMI adaptor molecule; generating a set of PCR duplicates of the tagged DNA molecule by performing PCR; and creating a single strand consensus sequence from all of the PCR duplicates which arose from an individual molecule of single-stranded DNA, resulting in two single strand consensus sequences.
30 . The method of claim 29 , further comprising a second stage of duplex consensus sequencing (DCS) which comprises
comparing the sequence of the two single strand consensus sequences arising from a single duplex DNA molecule; and reducing sequencing or PCR errors by (i) grouping the sequenced SMI-target nucleic acid products into families of paired target nucleic acid strands based on a common set of SMI sequences; and (ii) removing paired target nucleic acid strands having one or more nucleotide positions where the paired target nucleic acid strands disagree, or alternatively removing nucleotide positions from nucleic acid strands where the paired strands disagree at that specific position.
31 . The method of claim 29 , wherein the SMI adaptor molecules comprises a single molecule identifier (SMI) sequence, the SMI sequence comprising a degenerate or semi-degenerate nucleic acid sequence; and
an SMI ligation adaptor that allows the SMI adaptor molecule to be ligated to the double-stranded target nucleic acid sequence.
32 . The method of claim 31 , wherein the SMI sequence is single-stranded.
33 . The method of claim 31 , wherein the SMI sequence is double-stranded.
34 . The method of claim 32 , wherein the degenerate or semi-degenerate nucleic acid sequence comprises a degenerate nucleotide n-mer sequence.
35 . The method of claim 33 , wherein the double-stranded degenerate or semi-degenerate nucleic acid sequence comprises a first nucleotide n-mer sequence and a second n-mer sequence that is complementary to the first nucleotide n-mer sequence.
36 . The method of claim 35 , wherein the first nucleotide n-mer sequence is a degenerate sequence.Join the waitlist — get patent alerts
Track US2022195523A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.