US2023044849A1PendingUtilityA1

Using cell-free dna fragment size to determine copy number variations

Assignee: VERINATA HEALTH INCPriority: Feb 3, 2016Filed: Jul 22, 2022Published: Feb 9, 2023
Est. expiryFeb 3, 2036(~9.5 yrs left)· nominal 20-yr term from priority
G16B 25/00G16Z 99/00C12Q 1/6883G16B 20/00G06F 18/24323C12Q 1/6869G16B 30/10G16H 10/40G16H 50/30G16H 50/20G16H 20/10G16B 40/00C12Q 2600/154C12Q 1/68G16B 30/00C12Q 2600/156G16B 40/10C12Q 2537/16G16B 20/20C12Q 2537/165G16B 20/10G06F 18/23
77
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed are methods for determining copy number variation (CNV) known or suspected to be associated with a variety of medical conditions. In some embodiments, methods are provided for determining copy number variation of fetuses using maternal samples comprising maternal and fetal cell free DNA. In some embodiments, methods are provided for determining CNVs known or suspected to be associated with a variety of medical conditions. Some embodiments disclosed herein provide methods to improve the sensitivity and/or specificity of sequence data analysis by deriving a fragment size parameter. In some implementations, information from fragments of different sizes are used to evaluate copy number variations. In some implementations, one or more t-statistics obtained from coverage information of the sequence of interest is used to evaluate copy number variations. In some implementations, one or more fetal fraction estimates are combined with one or more t-statistics to determine copy number variations.

Claims

exact text as granted — not AI-modified
1 . A method for determining a copy number variation (CNV) of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the method comprising:
 (a) receiving sequence reads obtained by sequencing the cell-free nucleic acid fragments in the test sample;   (b) aligning the sequence reads of the cell-free nucleic acid fragments or aligning fragments containing the sequence reads to bins of a reference genome comprising the sequence of interest, thereby providing test sequence tags, wherein the reference genome is divided into a plurality of bins;   (c) determining fragment sizes of at least some of the cell-free nucleic acid fragments present in the test sample;   (d) determining coverages of the sequence tags for the bins of the reference genome;   (e) determining a t-statistic for the sequence of interest using coverages of bins in the sequence of interest and coverages of bins in a reference region for the sequence of interest; and   (f) determining a copy number variation in the sequence of interest in the test sample using a likelihood ratio calculated from the t-statistic and information about the sizes of the cell-free nucleic acid fragments.   
     
     
         2 . The method of  claim 1 , comprising performing (d) and (e) twice, once for fragments in a first size domain and again for fragments in a second size domain. 
     
     
         3 . The method of  claim 2 , wherein the first size domain comprises cell-free nucleic acid fragments of substantially all sizes in the sample, and the second size domain comprises only cell-free nucleic acid fragments smaller than a defined size. 
     
     
         4 . The method of  claim 2 , wherein the second size domain comprises only the cell-free nucleic acid fragments smaller than about 150 bp. 
     
     
         5 . The method of  claim 2 , wherein the likelihood ratio is calculated from a first t-statistic for the sequence of interest using sequence tags for fragments in a first size range, and a second t-statistic for the sequence of interest using sequence tags for fragments in a second size range. 
     
     
         6 . The method of  claim 1 , wherein the likelihood ratio is calculated as a first likelihood that the test sample is an aneuploid sample over a second likelihood that the test sample is a euploid sample. 
     
     
         7 . The method of  claim 1 , wherein the likelihood ratio is calculated from one or more values of fetal fraction in addition to the t-statistic and information about the sizes of the cell-free nucleic acid fragments. 
     
     
         8 . The method of  claim 7 , wherein the likelihood ratio is calculated from a fetal fraction, a t-statistic of short fragments, and a t statistics of all fragments, wherein the short fragments are cell-free nucleic acid fragments in a first size range smaller than a criterion size, and the all fragments are cell-free nucleic acid fragments including the short fragments and fragments longer than the criterion size. 
     
     
         9 . The method of  claim 8 , wherein the likelihood ratio is calculated:
     LR=Σ   ff     total     q ( ff   total )= p   1 ( T   short   ,T   all   |ff   est )/ p   0 ( T   short   ,T   all )   
       where p 1  represents the likelihood that data come from a multivariate normal distribution representing a 3-copy or 1-copy model, p 0  represents the likelihood that data come from a multivariate normal distribution representing a 2-copy model, T short , T all  are T scores calculated from chromosomal coverage generated from short fragments and all fragments, and q(ff total ) is a density distribution of the fetal fraction. 
     
     
         10 . The method of  claim 1 , further comprising:
 calculating values of a size parameter for the bins, wherein the size parameter comprises a value related to fragment size; and   determining a size-based t-statistic for the sequence of interest using values of the size parameter of bins in the sequence of interest and values of the size parameter of bins in the reference region for the sequence of interest.   
     
     
         11 . The method of  claim 10 , wherein the size parameter comprises a coverage weighted by fragment size. 
     
     
         12 . The method of  claim 10 , wherein the size parameter comprises a fraction of fragments in a defined size range. 
     
     
         13 . The method of  claim 10 , wherein the size parameter is biased toward a fragment size or size range. 
     
     
         14 . The method of  claim 10 , wherein the likelihood ratio of (f) is calculated from the t-statistic and the size-based t-statistic. 
     
     
         15 . The method of  claim 10 , wherein the likelihood ratio of (f) is calculated from the size-based t-statistic and a fetal fraction. 
     
     
         16 . The method of  claim 1 , further comprising comparing the likelihood ratio to a call criterion to determine a copy number variation in the sequence of interest. 
     
     
         17 . The method of  claim 16 , wherein the likelihood ratio is converted to a log likelihood ratio before being compared to the call criterion. 
     
     
         18 . The method of  claim 16 , wherein the call criterion is obtained by applying different criteria to a training set of training samples, and selecting a criterion that provides a defined sensitivity and a defined selectivity. 
     
     
         19 . The method of  claim 1 , further comprising obtaining a plurality of likelihood ratios and applying the plurality of likelihood ratios to a decision tree to determine a ploidy case for the sample. 
     
     
         20 . The method of  claim 1 , further comprising obtaining a plurality of likelihood ratios and one or more coverage values of the sequence of interest, and applying the plurality of likelihood ratios and one or more coverage values of the sequence of interest to a decision tree to determine a ploidy case for the sample. 
     
     
         21 . A system for evaluation of copy number of a nucleic acid sequence of interest in a test sample, the system comprising:
 a sequencer for receiving cell-free nucleic acid fragments from the test sample and providing nucleic acid sequence information of the test sample;   a processor; and   one or more computer-readable storage media having stored thereon instructions for execution on said processor to:
 (a) receive sequence reads obtained by sequencing the cell-free nucleic acid fragments in the test sample; 
 (b) align the sequence reads of the cell-free nucleic acid fragments or aligning fragments containing the sequence reads to bins of a reference genome comprising the sequence of interest, thereby providing test sequence tags, wherein the reference genome is divided into a plurality of bins; 
 (c) determine fragment sizes of at least some of the cell-free nucleic acid fragments present in the test sample; 
 (d) determine coverages of the sequence tags for the bins of the reference genome; 
 (e) determine a t-statistic for the sequence of interest using coverages of bins in the sequence of interest and coverages of bins in a reference region for the sequence of interest; and 
 (f) determine a copy number variation in the sequence of interest using a likelihood ratio calculated from the t-statistic and information about the sizes of the cell-free nucleic acid fragments. 
   
     
     
         22 . A method for determining a copy number variation (CNV) of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the method comprising:
 (a) receiving sequence reads obtained by sequencing the cell-free nucleic acid fragments in the test sample;   (b) aligning the sequence reads of the cell-free nucleic acid fragments or aligning fragments containing the sequence reads to bins of a reference genome comprising the sequence of interest, thereby providing test sequence tags, wherein the reference genome is divided into a plurality of bins;   (c) calculating coverages of the sequence tags for the bins of the reference genome;   (d) determining a t-statistic for the sequence of interest using coverages of bins in the sequence of interest and coverages of bins in a reference region for the sequence of interest;   (e) estimating one or more fetal fraction values of the cell-free nucleic acid fragments in the test sample; and   (f) determining a copy number variation in the sequence of interest using the t-statistic and the one or more fetal fraction values.   
     
     
         23 . A method for determining a copy number variation (CNV) of a nucleic acid sequence of interest in a test sample comprising cell-free nucleic acid fragments originating from two or more genomes, the method comprising:
 (a) receiving sequence reads obtained by sequencing the cell-free nucleic acid fragments in the test sample;   (b) aligning the sequence reads of the cell-free nucleic acid fragments or aligning fragments containing the sequence reads to bins of a reference genome comprising the sequence of interest, thereby providing test sequence tags, wherein the reference genome is divided into a plurality of bins;   (c) determining fragment sizes of the cell-free nucleic acid fragments existing in the test sample;   (d) calculating coverages of the sequence tags for the bins of the reference genome using sequence tags for the cell-free nucleic acid fragments having sizes in a first size domain;   (e) calculating coverages of the sequence tags for the bins of the reference genome using sequence tags for the cell-free nucleic acid fragments having sizes in a second size domain, wherein the second size domain is different from the first size domain;   (f) calculating size parameters for the bins of the reference genome using the fragment sizes determined in (c), wherein each size parameter comprises a value related to fragment size; and   (g) determining a copy number variation in the sequence of interest using the coverages calculated in (d) and (e) and the size parameters calculated in (f).   
     
     
         24 . The method of  claim 23 , wherein the first size domain comprises cell-free nucleic acid fragments of substantially all sizes in the sample, and the second size domain comprises only cell-free nucleic acid fragments smaller than a defined size. 
     
     
         25 . The method of  claim 24 , wherein the second size domain comprises only the cell-free nucleic acid fragments smaller than about 150 bp. 
     
     
         26 . The method of  claim 23 , wherein the size parameter comprises a coverage weighted by fragment size. 
     
     
         27 . The method of  claim 23 , wherein the size parameter comprises a fraction of fragments in a defined size range. 
     
     
         28 . The method of  claim 23 , wherein the size parameter is biased toward a fragment size or size range. 
     
     
         29 . A system for evaluation of copy number of a nucleic acid sequence of interest in a test sample, the system comprising:
 a sequencer for receiving cell-free nucleic acid fragments from the test sample and providing nucleic acid sequence information of the test sample;   a processor; and   one or more computer-readable storage media having stored thereon instructions for execution on said processor to:
 (a) receiving sequence reads obtained by sequencing the cell-free nucleic acid fragments in the test sample; 
 (b) aligning the sequence reads of the cell-free nucleic acid fragments or aligning fragments containing the sequence reads to bins of a reference genome comprising the sequence of interest, thereby providing test sequence tags, wherein the reference genome is divided into a plurality of bins; 
 (c) determining fragment sizes of the cell-free nucleic acid fragments existing in the test sample; 
 (d) calculating coverages of the sequence tags for the bins of the reference genome using sequence tags for the cell-free nucleic acid fragments having sizes in a first size domain; 
 (e) calculating coverages of the sequence tags for the bins of the reference genome using sequence tags for the cell-free nucleic acid fragments having sizes in a second size domain, wherein the second size domain is different from the first size domain; 
 (f) calculating size parameters for the bins of the reference genome using the fragment sizes determined in (c); and 
 (g) determining a copy number variation in the sequence of interest using the coverages calculated in (d) and (e) and the size parameters calculated in (f).

Join the waitlist — get patent alerts

Track US2023044849A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.