Using cell-free dna fragment size to determine copy number variations
Abstract
Disclosed are methods for determining copy number variation (CNV) associated with a variety of medical conditions. In some embodiments, methods are provided for determining copy number variation (CNV) of fetuses using maternal samples comprising maternal and fetal cell free DNA. In some embodiments, methods are provided for determining CNVs associated with a variety of medical conditions. Some embodiments disclosed herein provide methods to improve the sensitivity and/or specificity of sequence data analysis by deriving a fragment size parameter, such as a size-weighted coverage or a fraction of fragments in a size range. In some embodiments, the fragment size parameter is adjusted to remove within-sample GC-content bias. In some embodiments, removal of within-sample GC-content bias is based on sequence data corrected for systematic variation common across unaffected training samples. Also disclosed are systems and computer program products for evaluation of CNV of sequences of interest.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system for evaluation of copy number of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the system comprising one or more processors and system memory, the one or more processors being configured to:
align at least hundreds of thousands of sequence reads obtained from the cell-free nucleic acid fragments in the test sample to a reference genome comprising the nucleic acid sequence of interest, thereby providing at least hundreds of thousands of test sequence tags, wherein the reference genome is divided into a plurality of bins; obtain sizes of at least some of the cell-free nucleic acid fragments in the test sample; weight the at least hundreds of thousands of test sequence tags using one or more weight values or one or more weight functions, wherein the one or more weight values or weight functions are based on sizes of cell-free nucleic acid fragments from which the at least hundreds of thousands of test sequence tags are obtained; calculate coverages for the plurality of bins based on the weighted test sequence tags; and identify a copy number variation in the nucleic acid sequence of interest from the calculated coverages.
22 . The system of claim 21 , wherein the one or more weight values or one or more weight functions bias the coverages toward test sequence tags obtained from cell-free nucleic acid fragments in a size range characteristic of one genome in the test sample.
23 . The system of claim 22 , wherein the one or more weight functions comprise a non-linear function.
24 . The system of claim 23 , wherein the non-linear function comprise a Heaviside step function.
25 . The system of claim 23 , wherein the non-linear function comprise a box-car function.
26 . The system of claim 23 , wherein the non-linear function comprise a stair-case function.
27 . The system of claim 23 , wherein the non-linear function comprise a sigmoidal function.
28 . The system of claim 22 , wherein the one or more weight functions comprise a linear function.
29 . The system of claim 22 , wherein to weight the test sequence tags comprises to exclude test sequence tags obtained from cell-free nucleic acid fragments outside the size range.
30 . The system of claim 29 , wherein to weight the test sequence tags comprises to exclude fragments of a size greater than about 150 base pairs.
31 . The system of claim 21 , wherein to weight the at least hundreds of thousands of test sequence tags comprises to modify counts of test sequence tags by multiplication or exponentiation.
32 . The system of claim 21 , further comprising a sequencer for receiving cell-free nucleic acid fragments from the test sample and providing the sequence reads from the test sample.
33 . The system of claim 21 , wherein the two or more genomes comprise genomes from a mother and a fetus.
34 . The system of claim 33 , wherein the copy number variation in the nucleic acid sequence of interest comprises aneuploidy in the genome of the fetus.
35 . The system of claim 21 , wherein the two or more genomes comprise genomes from cancer and somatic cells.
36 . The system of claim 21 , wherein the copy number variation is related to a genetic abnormality.
37 . A method for determining a copy number variation of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the method comprising:
aligning at least hundreds of thousands of sequence reads of the cell-free nucleic acid fragments in the test sample to a reference genome comprising the nucleic acid sequence of interest, thereby providing at least hundreds of thousands of test sequence tags, wherein the reference genome is divided into a plurality of bins; determining sizes of at least some of the cell-free nucleic acid fragments in the test sample; weighting the at least hundreds of thousands of test sequence tags using one or more weight values or one or more weight functions, wherein the one or more weight values or weight functions are based on sizes of cell-free nucleic acid fragments from which the at least hundreds of thousands of test sequence tags are obtained; calculating coverages for the bins based on the weighted test sequence tags; and identifying a copy number variation in the nucleic acid sequence of interest from the calculated coverages.
38 . The method of claim 37 , wherein the one or more weight values or one or more weight functions bias the coverages toward test sequence tags obtained from cell-free nucleic acid fragments in a size range characteristic of one genome in the test sample.
39 . The method of claim 38 , wherein the one or more weight functions comprises a non-linear function.
40 . A computer program product comprising one or more computer-readable non-transitory storage media having stored thereon program code that, when executed by one or more processors of a computer system, cause the computer system to implement a method for determining a copy number variation of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the program code comprising code for:
aligning at least hundreds of thousands of sequence reads of the cell-free nucleic acid fragments in the test sample to a reference genome comprising the nucleic acid sequence of interest, thereby providing at least hundreds of thousands of test sequence tags, wherein the reference genome is divided into a plurality of bins; determining sizes of at least some of the cell-free nucleic acid fragments in the test sample; weighting the at least hundreds of thousands of test sequence tags using one or more weight values or one or more weight functions, wherein the one or more weight values or weight functions are based on sizes of cell-free nucleic acid fragments from which the at least hundreds of thousands of test sequence tags are obtained; calculating coverages for the bins based on the weighted test sequence tags; and identifying a copy number variation in the nucleic acid sequence of interest from the calculated coverages.Join the waitlist — get patent alerts
Track US2021371907A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.