US2023151417A1PendingUtilityA1
Library preparation and use thereof for sequencing-based error correction and/or variant identification
Est. expiryMar 31, 2037(~10.7 yrs left)· nominal 20-yr term from priority
C12N 15/1065C12Q 1/6855C12Q 1/6869G16B 30/00G16B 30/10C12Q 1/6806G16B 40/00
72
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Aspects of the invention include methods for preparing sequencing libraries, performing sequencing procedures that can correct for process-related errors, and identifying rare variants that are or may be indicative of cancer.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for preparing a sequencing library, the method comprising:
(a) obtaining a test sample comprising a plurality of double-stranded DNA (dsDNA) fragments; (b) providing a set of loop-shaped DNA adapters, wherein the set of loop-shaped DNA adapters comprises:
a plurality of first loop-shaped DNA adapters where each of the first loop-shaped DNA adapters comprises a single DNA molecule comprising (A) two complementary regions that hybridize with one another and leave an unpaired loop at the end of the DNA molecule, (B) an endonuclease restriction site in the unpaired loop, and (C) a single first unique molecular identifier (UMI),
a plurality of second loop-shaped DNA adapters where each of the first loop-shaped DNA adapters comprises a single DNA molecule comprising (A) two complementary regions that hybridize with one another and leave an unpaired loop at the end of the DNA molecule and (B) a single second unique molecular identifier (UMI);
(c) ligating the plurality of first loop-shaped DNA adapters to a first end of the dsDNA fragments and ligating the plurality of second loop-shaped adapters to a second end of the dsDNA fragments to generate a plurality of circular-shaped constructs, wherein each circular-shaped construct comprises a first loop-shaped DNA adapter ligated to a first end of the dsDNA fragment and a second loop-shaped DNA adapter ligated to a second end of the dsDNA fragment; and (d) after step (c), cleaving the plurality of first loop-shaped DNA adapters with an endonuclease to produce a plurality of linear single-strand DNA (ssDNA) molecules, wherein said linear ssDNA molecules comprise the forward strand and reverse complement strand.
2 . The method according to claim 1 , wherein the dsDNA fragments are cell-free DNA (cfDNA) fragments.
3 . The method according to claim 2 , wherein the cfDNA fragments originate from healthy cells and from cancer cells.
4 . The method according to claim 1 , wherein the test sample comprises whole blood, a blood fraction, plasma. serum, urine, feces, saliva, a tissue biopsy, pleural fluid, pericardial fluid, cerebral spinal fluid, peritoneal fluid, or any combination thereof.
5 . The method according to claim 1 , further comprising, between steps (a) and (b), modifying the plurality of dsDNA fragments by performing an end-repairing procedure and an A-tailing procedure prior to ligating the loop-shaped adapters to the dsDNA fragments.
6 . The method according to claim 1 , wherein the loop-shaped adapters further comprise a sample-specific index sequence.
7 . The method according to claim 1 , wherein the loop-shaped adapters further comprise a universal priming site.
8 . The method according to claim 1 , wherein the loop-shaped adapters further comprise one or more sequencing oligonucleotides for use in a cluster generation and/or a sequencing procedure.
9 . The method according to claim 1 , further comprising amplifying the linear ssDNA molecules.
10 . A method for generating a plurality of sequence reads for a plurality of the linear ssDNA molecules in a sequencing library, the method comprising:
(a) preparing a sequencing library according to the method of claim 1 ; and (b) sequencing a plurality of the linear ssDNA molecules in the sequencing library to generate a plurality of sequence reads.
11 . A method for correcting sequencing derived errors in sequence reads, the method comprising:
(a) generating a plurality of sequence reads according to the method of claim 10 ; (b) grouping the plurality of sequence reads into a plurality of families based on the first UMI and the second UMI, such that one or more unique nucleic acid sequence fragments originating from the same test sample contains the first UMI and the second UMI;
(i) including in the plurality of families each of the sequence reads that comprises both the first UMI and the second UMI; and
(ii) excluding from the plurality of families sequence reads that only comprise the first UMI on both ends of a dsDNA fragment or the second UMI on both ends of the dsDNA fragment;
(c) comparing the forward strand and the reverse complement strand of each of the sequence reads within each family of the plurality of families to generate a consensus sequence for each family, wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each nucleotide base is identified at a given position in the consensus sequence when a specific nucleotide base is present in at least 70% of the sequence reads of family members within each family of the plurality of families; and (d) using the consensus sequence from each of the plurality of families for error correction.
12 . The method according to claim 11 , wherein the dsDNA fragments are cell-free DNA (cfDNA) fragments.
13 . The method according to claim 12 , wherein the cfDNA fragments originate from healthy cells and from cancer cells.
14 . The method according to claim 11 , wherein the test sample comprises whole blood, a blood fraction, plasma, serum, urine, feces, saliva, a tissue biopsy, pleural fluid, pericardial fluid, cerebral spinal fluid, peritoneal fluid, or any combination thereof.
15 . The method according to claim 11 , wherein the consensus sequence for each family comprises a sequence of nucleotide bases, wherein each nucleotide base is identified at a given position in the consensus sequence when a specific nucleotide base is present in at least 80% of the sequence reads of family members within each family of the plural of families.
16 . The method according to claim 11 , wherein the consensus sequence for each family comprises a sequence of nucleotide bases, wherein each nucleotide base is identified at a given position in the consensus sequence when a specific nucleotide base is present in at least 90% of the sequence reads of family members within each family of the plurality of families.
17 . The method according to claim 11 , wherein the consensus sequence for each family comprises a sequence of nucleotide bases, wherein each nucleotide base is identified at a given position in the consensus sequence when a specific nucleotide base is present in at least 95% of the sequence reads of family members within each family of the plurality of families.
18 . The method according to claim 11 , further comprising loading at least a portion of the sequencing library into a sequencing flow cell and generating a plurality of sequencing clusters on the flow cell, wherein each of the sequencing clusters comprises the forward strand and the reverse complement strand.
19 . The method according to claim 11 , wherein the step of sequencing a plurality of the linear ssDNA molecules in the sequencing library comprises sequencing by a next-generation sequencing (NGS) procedure.
20 . The method according to claim 19 , wherein the NGS procedure comprises single-molecule real-time sequencing.
21 . The method according to claim 11 , wherein the step of sequencing a plurality of the linear ssDNA molecules in the sequencing library comprises a sequencing-by-synthesis procedure.
22 . The method according to claim 11 , wherein the step of sequencing a plurality of the linear ssDNA molecules in the sequencing library comprises a paired-end sequencing procedure.
23 . The method according to claim 11 , wherein the step of sequencing a plurality of the linear ssDNA molecules in the sequencing library comprises a single molecule sequencing procedure.
24 . A method for detecting one or more rare variants in a test sample, the method comprising:
(a) generating a plurality of sequence reads according to the method of claim 10 ; (b) grouping the plurality sequence reads into a plurality of families based on the first UMI and the second UMI, such that one or more unique nucleic acid sequence fragments originating from the same test sample contains the first UMI and the second UMI;
(i) including in the plurality of families each of the sequence reads that comprises both the first UMI and the second UMI; and
excluding from the plurality of families sequence reads that only comprise the first UMI on both ends of a dsDNA fragment or the second UMI on both ends of the dsDNA fragment;
(c) comparing the forward strand and the reverse complement strand of each of the sequence reads within each family of the plurality of families to a consensus sequence for each family; (d) aligning the consensus sequence for each family of the plurality of families to a reference sequence; and (e) identifying a consensus sequence as comprising a rare variant if the consensus sequence differs from the reference sequence at one or more nucleotide positions.
25 . The method according to claim 24 , wherein the dsDNA fragments are cell-free DNA (cfDNA) fragments.
26 . The method according to claim 25 , wherein the cfDNA fragments originate from healthy cells and from cancer cells.
27 . The method according to claim 24 , wherein the test sample comprises whole blood, a blood fraction, plasma, serum, urine, feces, saliva, a tissue biopsy, pleural fluid, pericardial fluid, cerebral spinal fluid, peritoneal fluid, or any combination thereof.
28 . The method according to claim 24 , further comprising using the one or more rare variants to detect a presence or absence of cancer, determine cancer status, monitor cancer progression, and/or determine a cancer classification.
29 . The method according to claim 28 , wherein the cancer comprises a carcinoma, a sarcoma, a myeloma, a leukemia, a lymphoma, a blastoma, a germ cell tumor, or any combination thereof.
30 . The method according to claim 24 , further comprising using the one or more rare variants to monitor cancer progression by monitoring disease progression, monitoring therapy, or monitoring cancer growth.
31 . The method according to claim 24 , further comprising using the one or more rare variants to determine cancer classification by determining cancer type and/or cancer tissue of origin.Join the waitlist — get patent alerts
Track US2023151417A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.