Methods for detecting copy-number variations in next-generation sequencing
Abstract
Copy Number Variants (CNV) detection methods described herein may efficiently integrate CNV detection into the workflow for a next generation sequencer (NGS) data processing, in parallel with SNP and INDEL variant calling. CNV detection methods as described herein may be performed by analyzing the coverage pattern across a suitable set of genomic regions or amplicons and across a batch of samples from different patients. The proposed methods do not require the use of specifically chosen reference samples as inputs to the workflow, but rather automatically select a set of reference samples from the same batch, for each sample being tested. The CNV detection methods may reliably detect CNVs in a set of samples without prior assumptions about the CNV status of any of those samples. Embodiments described herein may also apply the CNV detection scheme iteratively to further improve the detection performance, especially in the case of more frequent CNV occurrence. Since the knowledge on the CNVs in reference samples may improve their comparison with the sample being tested, the proposed methods may further comprise the step of iteratively feeding back the information about the CNVs found in the samples from any detection step into the next iteration step. The proposed methods may also further use additional information available from the NGS workflow about the samples, such as information on SNP fractions, as input to the NGS CNV detection.
Claims
exact text as granted — not AI-modified1 - 9 . (canceled)
10 . A method for detecting copy-number values (CNV) from a pool of DNA samples enriched with a target enrichment technology, each enriched DNA sample being associated with a library of pooled fragments from a set of amplicons/regions, each amplicon/region being sequenced with a high-throughput sequencer to generate coverage count for each sample and for each amplicon/region, comprising:
normalizing, with a data processing unit, the coverage count associated with each sample; selecting, with a data processing unit, for each sample, a set of reference samples as the samples with the closest normalized coverage count to the normalized coverage count of said sample, the number of reference samples in each subset of reference samples being a function of the total number of samples and being smaller than the total number of samples; for each sample, estimating the copy-number values in said sample as a function of at least the coverage counts in said sample and of at least the coverage counts in the selected set of reference samples for said sample.
11 . The method of claim 10 , wherein the number of reference samples N R in each set of reference samples is given by N R =[0.25*N]+2, where Nis the total number of samples.
12 . The method of claim 10 , wherein selecting a set of reference samples comprises calculating a distance between the coverage counts normalized both within each sample/plex and within each amplicon/region and selecting a set of samples with coverage counts having the shortest distances as the reference samples.
13 . The method of claim 12 , where the calculated distance is the Euclidean distance.
14 . The method of claim 10 , wherein the estimate of the copy-number values is calculated using information on the SNP fractions and coverage.
15 . The method of claim 10 , further comprising: applying a principal-component filter to the coverage count.
16 . The method of claim 10 , wherein normalizing the coverage count associated with each sample depends on a prior estimate of the copy number values for each sample and each amplicon/region.
17 . The method of claim 10 , wherein selecting a set of reference samples depends on a prior estimate of the copy number values for each sample and each amplicon/region.Join the waitlist — get patent alerts
Track US2022101944A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.