Methods and systems for generation and error-correction of unique molecular index sets with heterogeneous molecular lengths
Abstract
The disclosed embodiments concern methods, apparatus, systems and computer program products for determining sequences of interest using unique molecular index sequences that are uniquely associable with individual polynucleotide fragments, including sequences with low allele frequencies and long sequence length. In some implementations, the unique molecular index sequences include variable-length nonrandom sequences. In some implementations, the unique molecular index sequences are associated with the individual polynucleotide fragments based on alignment scores indicating similarity between the unique molecular index sequences and subsequences of sequence reads obtained from the individual polynucleotide fragments. System, apparatus, and computer program products are also provided for determining a sequence of interest implementing the methods disclosed.
Claims
exact text as granted — not AI-modified1 - 46 . (canceled)
47 . A method for sequencing nucleic acid molecules from a sample, comprising
(a) applying adapters to DNA fragments in the sample to obtain DNA-adapter products, wherein each adapter comprises a unique molecular index (UMI) in a set of unique molecular indices (UMIs); (b) amplifying the DNA-adapter products to obtain a plurality of amplified polynucleotides; (c) sequencing the plurality of amplified polynucleotides, thereby obtaining a plurality of reads associated with the set of UMIs; (d) obtaining, for each read of the plurality of reads, alignment scores with respect to the set of UMIs, each alignment score indicating similarity between a subsequence of a read and a UMI; (e) identifying, among the plurality of reads, reads associated with a same UMI using the alignment scores; and (f) determining a sequence of a DNA fragment in the sample using the reads associated with the same UMI.
48 . The method of claim 47 , wherein the alignment scores are based on matches of nucleotides and edits of nucleotides between the subsequence of the read and the UMI.
49 . The method of claim 48 , wherein the edits of nucleotides comprise substitutions, additions, and deletions of nucleotides.
50 . The method of claim 48 , wherein each alignment score penalizes mismatches at the beginning of a sequence but does not penalize mismatches at the end of the sequence.
51 . The method of claim 50 , wherein obtaining an alignment score between a read and a UMI comprises:
(a) calculating an alignment score between the UMI and each one of all possible prefix sequences of the subsequence of the read; (b) calculating an alignment score between the subsequence of the read and each one of all possible prefix sequences of the UMI; and (c) obtaining a largest alignment score among the alignment scores calculated in (a) and (b) as the alignment score between the read and the UMI.
52 . The method of claim 47 , wherein the set of UMIs comprises UMIs of at least two different molecular lengths.
53 . The method of claim 47 , wherein the subsequence has a length that equals to a length of the longest UMI in the set of UMIs.
54 . The method of claim 47 , wherein identifying the reads associated with the same UMI in (e) further comprises:
selecting, for each read of the plurality of reads, at least one UMI from the set of UMIs based on the alignment scores; and associating each read of the plurality of reads with the at least one UMI selected for the read.
55 . The method of claim 54 , wherein selecting the at least one UMI from the set of UMIs comprises selecting a UMI having a highest alignment score among the set of UMIs.
56 . The method of claim 54 , wherein the at least one UMI comprises two or more UMIs.
57 . The method of claim 56 , further comprising selecting one of the two or more UMI as the same UMI of (e) and (f).
58 . The method of claim 47 , wherein the set of UMIs consist of variable-length, nonrandom UMIs (vNRUMIs).
59 . The method of claim 47 , wherein (a) further comprises applying adapters to both ends of the DNA fragments in the sample.
60 . The method of claim 47 , wherein (e) comprises collapsing reads associated with the same UMI into a group to obtain a consensus nucleotide sequence for the sequence of the DNA fragment in the sample.
61 . The method of claim 60 , the consensus nucleotide sequence is obtained based partly on quality scores of the reads.
62 . The method of claim 47 , wherein (f) comprises:
identifying, among the reads associated with the same UMI, reads having a same read position or similar read positions in a reference sequence, and determining the sequence of the DNA fragment using reads that (i) are associated with the same UMI and (ii) have the same read position or similar read positions in the reference sequence.
63 . The method of claim 47 , wherein the set of UMIs includes no more than about 10,000 different UMIs.
64 . The method of claim 63 , wherein the set of UMIs includes no more than about 1,000 different UMIs.
65 . A computer program product comprising a non-transitory machine readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for sequencing nucleic acid molecules from a sample, said program code comprising:
(a) code for obtaining a plurality of reads of a plurality of amplified polynucleotides, each polynucleotide of the plurality of amplified polynucleotides comprising an adapter attached to a DNA fragment, wherein the adapter comprises a unique molecular index (UMI) in a set of unique molecular indices (UMIs); (b) code for obtaining, for each read of the plurality of reads, alignment scores with respect to the set of UMIs, each alignment score indicating similarity between a subsequence of a read and a UMI; (c) code for identifying, among the plurality of reads, reads associated with a same UMI based in part on the alignment score; and (d) code for determining, using the reads associated with the same UMI, a sequence of a DNA fragment in the sample.
66 . A computer system, comprising:
one or more processors; system memory; and one or more computer-readable storage media having stored thereon computer-executable instructions that causes the computer system to implement a method for determine sequence information of a sequence of interest in a sample, the instructions comprising: (a) obtaining a plurality of reads of a plurality of amplified polynucleotides, each polynucleotide of the plurality of amplified polynucleotides comprising an adapter attached to a DNA fragment, wherein the adapter comprises a unique molecular index (UMI) in a set of unique molecular indices (UMIs); (b) obtaining, for each read of the plurality of reads, alignment scores with respect to the set of UMIs, each alignment score indicating similarity between a subsequence of a read and a UMI; (c) identifying, among the plurality of reads, reads associated with a same UMI based in part on the alignment score; and (d) determining, using the reads associated with the same UMI, a sequence of a DNA fragment in the sample.Join the waitlist — get patent alerts
Track US2024011087A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.