Sequencing Process
Abstract
The present invention relates to methods for generating sequences of template nucleic acid molecules, methods for determining sequences of at least two template nucleic acid molecules, computer programs adapted to perform the methods and computer readable media storing the computer programs. In particular the present invention relates to methods for generating sequences of at least one individual target template nucleic acid molecule comprising: a) providing at least one sample of nucleic acid molecules comprising at least two target template nucleic acid molecules; b) introducing a first molecular tag into one end of each of the at least two target template nucleic acid molecules and a second molecular tag into the other end of each of the at least two target template nucleic acid molecules to provide at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules is tagged with a unique first molecular tag and a unique second molecular tag; c) amplifying the at least two tagged template nucleic acid molecules to provide multiple copies of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag; d) sequencing regions of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag; and e) reconstructing a consensus sequence for at least one of the at least two target template nucleic acid molecules.
Claims
exact text as granted — not AI-modified1 .- 30 . (canceled)
31 . A method for generating sequences of at east one individual target tern ate nucleic acid molecule comprising:
a) providing at least one sample of n is acid molecules comprising at least two target template nucleic acid molecules; b) introducing a first molecular tag into one end of each of the at least two target template nucleic acid molecules and a second molecular tag into the other end of each of the at least two target template nucleic acid molecules to provide at least two tagged template nucleic acid molecules wherein each tagged template nucleic acid molecule is tagged with a unique first molecular tag and a unique second molecular tag; c) amplifying the at least two tagged template nucleic acid molecules to provide multiple copies of the at least two tagged template nucleic acid molecules; d) sequencing regions of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag; and e) reconstructing a consensus sequence for at least one of the at least two target template nucleic acid molecules; wherein step e) comprises
(i) identifying clusters of sequences of the regions of the multiple copies of the at east two tagged template nucleic acid molecules which are likely to correspond to the same target template nucleic acid molecule by assigning sequences comprising first molecular tag sequences which are homologous to one another and second molecular tag sequences which are homologous to one another to the same cluster;
(ii) selecting at least one cluster of sequences wherein the sequences within the selected clusters comprise a first molecular tag and a second molecular tag which are more commonly associated with one another than with a different first molecular tag or second molecular tag;
(iii) reconstructing a consensus sequence of a first target template nucleic acid molecule by aligning sequences of the at least two template nucleic acid molecules in the cluster selected in step (ii) and defining a consensus sequence from these sequences; and
(iv) performing steps (ii) to (iii) in respect of a second and/or further template nucleic acid molecule.
32 . A method for generating sequences of at least one individual target template nucleic acid molecule which is greater than 1 Kbp in size comprising:
a) providing at least one sample of nucleic acid molecules comprising at least two target template nucleic acid molecules which are greater than 1 Kbp in size; b) introducing a first molecular tag into one end of each of the at least two target template nucleic acid molecules and a second molecular tag into the other end of each of the at least two target template nucleic acid molecules to provide at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules is tagged with a unique first molecular tag and a unique second molecular tag; c) amplifying the at least two tagged template nucleic acid molecules to provide multiple copies of the at least two tagged template nucleic acid molecules; d) isolating a fraction of the multiple copies of the at least two tagged template nucleic acid molecules and fragmenting the tagged template nucleic acid molecules in the fraction to provide multiple fragmented template nucleic acid molecules; e) sequencing regions of the multiple copies of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag; f) sequencing the multiple fragmented template nucleic acid molecules; and g) reconstructing a consensus sequence for at least one of the at least two target template nucleic acid molecules from sequences comprising at least a subset of the sequences produced in step f).
33 . The method of claim 32 , wherein
(A) the method further comprises a step of enriching the multiple fragmented template molecules to increase the proportion of the multiple fragmented template nucleic acid molecules comprising the first molecular tag or the second molecular tag and wherein this step is before step f); and/or (B) step g) comprises:
(i) identifying clusters of sequences of regions of the multiple copies of the at least two tagged template nucleic acid molecules which are likely to correspond to the same individual target template nucleic acid molecule by assigning sequences comprising first molecular tag sequences which are homologous to one another and second molecular tag sequences which are homologous to one another to the same cluster;
(ii) analysing the sequences of the multiple fragmented template nucleic acid molecules to identify sequences of the multiple fragmented template nucleic acid molecules which comprise a first molecular tag which is homologous to the first molecular tag of the sequences of a first cluster or a second molecular tag which is homologous to the second molecular tag of the sequences of the first cluster;
(iii) reconstructing the sequence of a first template nucleic acid molecule by aligning sequences comprising at least a subset of the sequences of the multiple fragmented template nucleic acid molecules identified in step (ii) and defining a consensus sequence from these sequences; and
(iv) performing steps (i) to (iii) in respect of a second and/or further template nucleic acid molecule.
34 . A computer-implemented method for determining sequences of at least one individual target template nucleic acid molecule comprising the following steps:
(a) obtaining data comprising sequences of regions of multiple copies of at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules comprises a first molecular tag at one end and a second molecular tag at the other end, wherein each target template nucleic acid molecule is tagged with a unique first molecular tag and a unique second molecular tag and wherein the regions comprise the first molecular tag and the second molecular tag; (b) analysing the data comprising sequences of regions of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag to identify clusters of sequences which are likely to correspond to the same individual target template nucleic acid molecule by assigning sequences comprising first molecular tags which are homologous to one another and second molecular tags which are homologous to one another to the same cluster; (c) obtaining data comprising sequences of multiple fragments of the at least two tagged template nucleic acid molecules wherein each of the fragments comprise either the first molecular tag or the second molecular tag; (d) analysing the sequences of the multiple fragments of the at least two tagged template nucleic acid molecules to identify sequences of the multiple fragments of the at least two tagged template nucleic acid molecules which comprise the first molecular tag which is homologous to the first molecular tag of the sequences of a first cluster or the second molecular tag which is homologous to the second molecular tag of the sequences of the first cluster; (e) reconstructing the sequence of a first target template nucleic acid molecule by aligning sequences comprising at least a subset of the sequences of the multiple fragments of the at least two tagged template nucleic acid molecules identified in step (d) and defining a consensus sequence from these sequences; and (f) performing steps (c) to (e) in respect of a second and/or further target template nucleic acid molecule.
35 . A computer-implemented method for determining sequences of at least one target template nucleic acid molecule comprising the following steps:
(a) obtaining data comprising clusters of sequences wherein:
(i) each cluster comprises sequences of regions of multiple copies of at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules comprises a first molecular tag at one end and a second molecular tag at the other end, wherein each target template nucleic acid is tagged with a unique first molecular tag and a unique second molecular tag and wherein the regions comprise the first molecular tag and the second molecular tag;
(ii) each cluster comprises sequences of multiple fragments of the at least two tagged template nucleic acid molecules wherein each of the fragments comprises either the first molecular tag or the second molecular tag;
(iii) the sequences of regions of multiple copies of at least two tagged template nucleic acid molecules in each cluster comprise first molecular tags and second molecular tags which are homologous to one another;
(iv) the sequences of the multiple fragments of the at least two tagged template nucleic acid molecules comprise the first molecular tag which is homologous to the first molecular tag of the sequences of regions of the multiple copies of at least two tagged template nucleic acid molecules in that cluster or the second molecular tag which is homologous to the second molecular tag of the sequences of regions of multiple copies of the at least two tagged template nucleic acid molecules in that cluster;
(b) reconstructing the sequence of a first target template nucleic acid molecule by aligning sequences comprising at least a subset of the sequences of the multiple fragments of the at least two tagged template nucleic acid molecules in a first cluster and defining a consensus sequence from these sequences; and (c) performing step (b) in respect of a second and/or further template nucleic acid molecule.
36 . The method of claim 33 , wherein step (i) further comprises determining a consensus sequence for the first molecular tag sequences and a consensus sequence for the second molecular tag sequences of a first cluster and step (ii) comprises identifying sequences of the multiple fragmented template nucleic acid molecules which comprise a first molecular tag or a second molecular tag which is homologous to the consensus sequence for the first molecular tag or the consensus sequence for the second molecular tag of the first cluster.
37 . The method of claim 34 , wherein step (b) further comprises determining a consensus sequence for the first molecular tag sequences and a consensus sequence for the second molecular tag sequences of a first cluster and step (d) comprises identifying sequences of the multiple fragmented template nucleic acid molecules which comprise a first molecular tag or a second molecular tag which is homologous to the consensus sequence for the first molecular tag or the consensus sequence for the second molecular tag of the first cluster.
38 . The method of claim 32 , further comprising steps of:
(v) identifying clusters of sequences of regions of the multiple copies of the at least two tagged template nucleic acid molecules which are likely to correspond to the same template nucleic acid molecule by assigning sequences comprising first molecular tag sequences which are homologous to one another and second molecular tag sequences which are homologous to one another to the same cluster; (vi) selecting at least one cluster of sequences wherein the sequences within the selected clusters comprise a first molecular tag and a second molecular tag which are more commonly associated with one another than with a different first molecular tag or second molecular tag;
wherein the sequence of the first target template nucleic acid molecule is reconstructed from the sequences in the cluster selected in step (vi), optionally
wherein step (vi) consists of identifying groups of clusters of sequences of the at least two tagged. template nucleic acid molecules wherein the sequences within the clusters of each group have first molecular tags which are homologous to one another and/or identifying groups of clusters of sequences of the at least two tagged template nucleic acid molecules wherein the sequences within the clusters of each group have second molecular tags which are homologous to one another and selecting a cluster from the group of clusters of sequences wherein the cluster that is selected contains the highest number of sequences.
39 . A computer-implemented method for determining sequences of at least one individual target template nucleic acid molecule comprising the following steps:
(a) obtaining data comprising sequences of regions of multiple copies of at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules comprises a first molecular tag at one end and a second molecular tag at the other end, wherein each target template nucleic acid molecules is tagged with a unique first molecular tag and a unique second molecular tag and wherein the regions comprise the first molecular tag and the second molecular tag; (b) analysing the data comprising sequences of regions of the at least two tagged template nucleic acid molecules comprising the first molecular tag and the second molecular tag to identify clusters of sequences which are likely to correspond to the same template nucleic acid molecule by assigning sequences comprising first molecular tags which are homologous to one another and second molecular tags which are homologous to one another to the same cluster; (c) selecting at least one cluster of sequences wherein the sequences within the selected clusters comprise a first molecular tag and a second molecular tag which are more commonly associated with one another than with a different first molecular tag or second molecular tag; (d) reconstructing a consensus sequence of a first target template nucleic acid molecule by aligning at least a subset of the sequences molecules in the cluster selected in step (c) and defining a consensus sequence from these sequences; and (e) performing steps (c) to (d) in respect of a second and/or further target template nucleic acid molecule.
40 . The method step of claim 31 , wherein (iv) consisting of identifying groups of clusters of sequences of the at least two tagged template nucleic acid molecules wherein the sequences within the clusters of each group have 5′ molecular tags which are homologous to one another and/or identifying groups of clusters of sequences of the at least two tagged template nucleic acid molecules wherein the sequences within the clusters of each group have 3′ molecular tags which are homologous to one another; and selecting a cluster from a group of clusters of sequences wherein the cluster that is selected contains the highest number of sequences.
41 . A computer-implemented method for determining sequences of at least one target template nucleic acid molecule comprising
(a) obtaining data comprising a cluster of sequences; (b) reconstructing a consensus sequence of a first template nucleic acid molecule by aligning the sequences of at least a subset of the sequences in the selected cluster; wherein the sequences in the selected cluster comprise sequences of regions of multiple copies of at least two tagged template nucleic acid molecules wherein each of the at least two tagged template nucleic acid molecules comprises a first molecular tag at one end and a second molecular tag at the other end, wherein each of the at least two target template nucleic acid molecules is tagged with a unique first molecular tag and a unique second molecular tag and wherein the regions comprise the first molecular tag and the second molecular tag; and each sequence in the selected cluster
(i) comprises first molecular tag which is homologous to the molecular tag of the other sequences in that cluster and the second molecular tag which is homologous to the second molecular tag of the other sequences in that cluster;
(ii) comprises a first molecular tag and a second molecular tag which are more commonly associated with one another than with a different first molecular tag or second molecular
42 . The method of claim 34 , wherein:
(A) the first molecular tags of the sequences of the same cluster have at least 90% sequence identity to one another; and/or (B) the second molecular tags of the sequences of the same cluster have at least 90% sequence identity to one another.
43 . The method of claim 32 wherein
(A) step g) is a computer-implemented method step; and/or
(B) which is a computer-implemented method; and/or
(C) steps e) and/or are carried out using sequencing technology comprising a step of bridge PCR; and/or
(D) steps e) and/or f) are carried out using sequencing technology comprising a step of bridge PCR and the step of bridge PCR is carried out using an extension time of greater than 15 seconds; and/or
(E) steps e) and f) are carried out in different sequencing runs.
44 . The method of claim 38 , wherein:
(A) step (e) is a computer-implemented method step; and/or (B) which is a computer-implemented method; and/or (C) step d) is carried out using sequencing technology comprising a step of bridge PCR.
45 . The method of claim 31 , wherein:
(A) the regions comprise greater than 25 base pairs comprising the first molecular tag or the second molecular tag; and/or (B) the regions comprise the entire length of the a least two tagged template nucleic acid molecules are sequenced: and/or (C) the first molecular tag and the second molecular tag are introduced into the at least two template nucleic acid molecules using a method selected from the group consisting of PCR, tagmentation, and physical shearing or restriction digestion of the at least one template nucleic acid molecule followed by ligation of nucleic acids comprising the 5′ molecular tag or the 3′ molecular tag.
46 . The method of claim 45 , wherein in (C) the first molecular tag and the second molecular tag are introduced into the at least two template nucleic acid molecules by PCR using primers comprising a portion comprising the first molecular tag or the second molecular tag and a portion having a sequence that is capable of hybridising to the at least two template nucleic acid molecules.
47 . The method of claim 31 , wherein:
(A) the at least two template nucleic acid molecules encode microbial ribosomal 16S; and/or (B) at least one of the at least two template nucleic acid molecules is less than 10 Kbp in size.
48 . A computer program adapted to perform the method of claim 34 , when said program is run on an electronic device.
49 . A computer readable medium storing the computer program of claim 48 .Join the waitlist — get patent alerts
Track US2021403991A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.