US2017240973A1PendingUtilityA1
Methods to determine tumor gene copy number by analysis of cell-free dna
Est. expiryDec 17, 2035(~9.4 yrs left)· nominal 20-yr term from priority
G16B 99/00C12Q 1/6886G16B 30/00C12Q 1/6874C12Q 1/6809G06F 19/22G16B 20/10G16B 30/10
61
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods are provided herein to improve automatic detection of copy number variation in nucleic acid samples. These methods provide improved approaches for determining baseline copy number of genetic loci within a sample, reduce variation due to features of genetic loci, sample preparation, and probe exhaustion.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
(a) obtaining sequencing reads of deoxyribonucleic acid (DNA) molecules of a cell-free bodily fluid sample of a subject; (b) generating from the sequencing reads a first data set comprising, for each genetic locus in a plurality of genetic loci, a quantitative measure related to sequencing read coverage (“read coverage”); (c) correcting the first data set by performing saturation equilibrium correction and probe efficiency correction; (d) determining a baseline read coverage for the first data set, wherein the baseline read coverage relates to saturation equilibrium and probe efficiency; and (e) determining a copy number state for each genetic locus in the plurality of genetic loci relative to the baseline read coverage.
2 . The method of claim 1 , wherein the first data set comprises, for each genetic locus in a plurality of genetic loci, a quantitative measure related to guanine-cytosine content of the genetic locus (“GC content”).
3 . The method of claim 2 , comprising prior to (c) removing from the first data set genetic loci that are high-variance genetic loci, wherein removing comprises:
(i) fitting a model relating the quantitative measures related to guanine-cytosine content and the quantitative measures of sequencing read coverage of the genetic loci; and (ii) removing from the first data set at least 10% of the plurality of genetic loci, wherein removing the genetic loci comprises removing the at least 10% of the plurality of genetic loci that most differ from the model, thereby providing the first data set of baselining genetic loci.
4 . (canceled)
5 . The method of claim 3 , wherein performing saturation equilibrium correction comprises transforming the first set data set of baselining genetic loci into a saturation corrected data set by:
(i) determining for each genetic locus from the first data set of baselining genetic loci a quantitative measure related to a probability that a strand of DNA molecule from the sample derived from the genetic locus is represented within the sequencing reads; (ii) determining a first transformation for the read coverage by relating the read coverage in the first data set of baselining genetic loci to both the GC content of the first data set of baselining genetic loci and the quantitative measure related to the probability that a strand of DNA derived from each genetic locus in the first data set of baselining genetic loci is represented within the sequencing reads; and (iii) applying the first transformation to the read coverage of each genetic locus from the first data set of baselining genetic loci to provide the saturation corrected data set, wherein the saturation corrected data set comprises a first set of transformed read coverages of the first data set of baselining genetic loci;
6 . The method of claim 5 , wherein determining the first transformation comprises (i) determining a measure related to central tendency of the read coverage of the first data set of baselining genetic loci; (ii) determining a function that fits the measure related to central tendency of the read coverage of the first data set of baselining genetic loci based on the GC content of the genetic locus and the quantitative measure related to the probability that a strand of DNA derived from the genetic locus is represented within the sequencing reads; and (iii) for each genetic locus of the first data set of baselining genetic loci, determining a difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage.
7 . (canceled)
8 . (canceled)
9 . The method of claim 5 , wherein performing probe efficiency correction comprises transforming the saturation corrected data set into a probe efficiency corrected data set by:
(i) removing from the saturation corrected data set genetic loci that are high-variance genetic loci with respect to the first set of transformed read coverages, thereby providing a second data set of baselining genetic loci; (ii) determining a second transformation for the first set of transformed read coverages related to the probe efficiency of the second data set of baselining genetic loci; and (iii) transforming the first set of transformed read coverages of the second data set of baselining genetic loci with the second transformation, thereby providing the probe efficiency corrected data set, wherein the probe efficiency corrected data set comprises a second set of transformed read coverages of the second data set of baselining genetic loci.
10 . The method of claim 9 , wherein removing from the first data set genetic loci that are high-variance genetic loci comprises:
(i) fitting a model relating the GC content and the first set of transformed read coverages of the saturation corrected data set; and (ii) removing from saturation corrected data set at least 10% of the genetic loci, wherein the removing the genetic loci comprises removing genetic loci that most differ from the model, thereby providing the second data set of baselining genetic loci.
11 . (canceled)
12 . (canceled)
13 . (canceled)
14 . (canceled)
15 . (canceled)
16 . The method of claim 5 , further comprising:
(g) determining a third transformation for the second set of transformed read coverages by relating the transformed read coverages of the second data set of baselining genetic loci to both the GC content of the second data set of baselining genetic loci and the quantitative measure related to the probability that a strand of DNA derived from the each locus in the second data set of baselining genetic loci is represented within the sequencing reads; (h) applying the third transformation to the second set of transformed read coverages to provide a fourth data set, wherein the fourth data set comprises a third set of transformed quantitative read coverages.
17 . The method of claim 1 , wherein the DNA molecules of the cell-free bodily fluid sample are enriched for the plurality of genetic loci using one or more oligonucleotide probes that are complementary to at least a portion of the genetic loci from the plurality of genetic loci.
18 . The method of claim 17 , wherein the GC content of each genetic locus from the plurality of genetic loci is a measure related to central tendency of guanine-cytosine content of the one or more oligonucleotide probes that are complementary to at least a portion of the genetic loci from the plurality of genetic loci.
19 . The method of claim 17 , wherein the read coverage of the genetic locus is a measure related to central tendency of the read coverage of regions of the genetic locus corresponding to the one or more oligonucleotide probes.
20 . The method of claim 17 , wherein the performing saturation equilibrium correction and the performing probe efficiency correction comprise fitting a Langmuir model, wherein the Langmuir model comprises probe efficiency (K) and saturation equilibrium constant (I sat ).
21 . (canceled)
22 . (canceled)
23 . (canceled)
24 . (canceled)
25 . (canceled)
26 . (canceled)
27 . (canceled)
28 . The method of claim 1 , wherein obtaining the sequencing reads comprises ligating adaptors to the DNA molecules from the cell-free bodily fluid sample from the subject.
29 . The method of claim 28 , wherein the DNA molecules comprise duplex DNA molecules and the adaptors are ligated to the duplex DNA molecules such that each adaptor differently tags complementary strands of the DNA molecule to provide tagged strands.
30 . The method of claim 29 , wherein determining the quantitative measure related to the probability that a strand of DNA derived from the genetic locus is represented within the sequencing reads comprises sorting sequencing reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads.
31 . The method of claim 30 , further comprising determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci to determine a quantitative measure related to total double-stranded DNA molecules in said sample that map to each of said one or more genetic loci based on said quantitative measure related to paired reads and unpaired reads mapping to each locus.
32 . The method of claim 28 , wherein the adaptors comprise barcode sequences.
33 . The method of claim 32 , wherein determining the read coverage comprises collapsing the sequencing reads based on position of the mapping of the sequencing reads to the reference genome and the barcode sequences.
34 . (canceled)
35 . The method of claim 1 , further comprising determining that at least a subset of the baselining genetic loci have undergone copy number alteration in the tumor cells of the subject by determining relative quantities of variants within the baselining genetic loci for which the germline genome of the subject is heterozygous.
36 . The method of claim 35 , wherein the relative quantities of the variants are not approximately equal, and wherein the baselining genetic loci for which the relative quantities of the variants are not approximately equal are removed from the baselining genetic loci, thereby providing allelic-frequency corrected baselining genetic loci.
37 . (canceled)
38 . (canceled)
39 . (canceled)
40 . (canceled)
41 . (canceled)Join the waitlist — get patent alerts
Track US2017240973A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.