US2015056619A1PendingUtilityA1

Method and system for determining copy number variation

Assignee: LI XUCHAOPriority: Apr 5, 2012Filed: Apr 5, 2012Published: Feb 26, 2015
Est. expiryApr 5, 2032(~5.7 yrs left)· nominal 20-yr term from priority
C12Q 1/6869G06F 19/22C12Q 1/68G16B 15/00G16B 30/10G16C 99/00C12Q 2537/165C12Q 2535/122G16B 40/00G16B 30/00
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed are a method and a system for determining genome copy number variation, which relates to the technical field of bioinformatics. The method comprises obtaining reads; determining sequence labels according to the reads; counting the number of sequence labels falling into each window; performing GC correction on the sequence label number of each window and a correction according to an expected sequence label number adjusted by a control set to obtain a corrected sequence label number; selecting a demarcation point with a small significance value as a candidate CNV breaking point; rejecting the least significant candidate CNV breaking point at every turn, updating difference significance values of two candidate CNV breaking points on the left and right of the rejected candidate CNV breaking point and performing cyclic iteration until difference significance values of all candidate CNV breaking points are smaller than a termination threshold value, thereby determining a CNV breaking point. The method and the system the present invention have clinical feasibility, and can precisely detect a micro-deletion/micro-duplication area of 0.5 M under the situation of using data of about 50 M.

Claims

exact text as granted — not AI-modified
1 . A method of detecting a copy number variation comprising following steps:
 obtaining reads from at least one part of a nucleic acid molecule of a sample,   determining uniquely-mapped reads aligned to a (genomic) reference sequence based on the obtained reads,   dividing the genomic reference sequence into a plurality of windows, and calculating the number of uniquely-mapped reads falling into each of the plurality of windows,   subjecting the number of uniquely-mapped reads falling into each of the plurality of windows to a GC correction, and to a correction based on an expected number of uniquely-mapped reads adjusted by a control set to obtain a corrected number of uniquely-mapped reads,   calculating a significance value of the difference between two numerical populations each consisting of the corrected numbers of uniquely-mapped reads falling into windows on each of the two sides of a demarcation point, the demarcation point being a starting point or an ending point of each of the plurality of windows, to thereby select the demarcation point having a smaller significance value as a candidate CNV breakpoint;   calculating a significance value of the difference between two numerical populations each consisting of the corrected number of uniquely-mapped reads falling into windows contained within each of two sequences, with one sequence ranging from a given candidate CNV breakpoint to an adjacent upstream candidate CNV breakpoint, and the other sequence ranging from the given candidate CNV breakpoint to an adjacent downstream candidate CNV breakpoint, and   removing the candidate CNV breakpoint having the least significance at every turn and recalculating the significance value for the two candidate CNV breakpoints adjacent to the removed candidate CNV breakpoint, performing cyclic iteration until the significance values of all candidate CNV breakpoints are less than a termination threshold value, to thereby determine the CNV breakpoint.   
     
     
         2 . The method of  claim 1 , further comprising a step of:
 subjecting the at least one part of the nucleic acid molecule of the sample to sequencing, to obtain the reads.   
     
     
         3 . The method of  claim 1 , wherein each of the plurality of windows contains the same number of reference unique reads, or each of the plurality of windows has the same length. 
     
     
         4 . The method of  claim 1 , wherein the termination threshold value is obtained based on the control set consisting of normal samples. 
     
     
         5 . The method of  claim 1 , wherein the step of subjecting the number of uniquely-mapped reads falling into each of the plurality of windows to the GC correction and to a correction based on the expected number of uniquely-mapped reads adjusted by the control set to obtain the corrected number of uniquely-mapped reads further comprises one or more selected from the group consisting of:
 grouping the plurality of windows based on a GC content, and obtaining a correction coefficient based on a mean value of the number of uniquely-mapped reads within one group and a mean value of the number of uniquely-mapped reads for all of the plurality of windows, and then subjecting the number of uniquely-mapped reads falling into each of the plurality of windows to the correction, to thereby obtain the GC-corrected number of uniquely-mapped reads; and   obtaining the expected number of uniquely-mapped reads adjusted by the control set by following steps:   calculating a ratio of the GC-corrected number of uniquely-mapped reads falling into each of the plurality of windows in the control set to the total number of uniquely-mapped reads;   obtaining a mean value of the ratios for all windows corresponding to the control set; and   calculating the expect number of uniquely-mapped reads for each of the plurality of windows in the sample based on obtained mean value of the ratios and the total number of uniquely-mapped reads in the sample.   
     
     
         6 . The method of  claim 5 , wherein a cyclization with a single chromosome or a whole genome is performed when selecting the candidate CNV breakpoint. 
     
     
         7 . The method of  claim 1 , wherein after the CNV breakpoint is determined, the method further comprises:
 subjecting a sequence between two CNV breakpoints to a confidence selection, wherein the confidence selection comprises steps of:   determining a confidence interval of the normal corrected number of uniquely-mapped reads using the control set, based on a distribution of the corrected number of uniquely-mapped reads; and   determining an abnormality presenting in the sequence between the two CNV breakpoints, if the mean value of the corrected number of uniquely-mapped reads within the sequence is outside the confidence interval.   
     
     
         8 . The method of  claim 7 , wherein the corrected number of uniquely-mapped reads fits a normal distribution, and the confidence interval is 95%. 
     
     
         9 . The method of  claim 1 , wherein the method further comprises one or more selected from the group consisting of:
 obtaining the sample from a human, the sample comprising amniotic fluid obtained by amniocentesis, villus obtained by chorionic villi sampling, umbilical cord blood obtained by percutaneous umbilical blood sampling, spontaneous miscarrying fetus tissue or human peripheral blood;   obtaining a genomic DNA of the sample by a DNA extraction method such as a salting-out method, a column chromatography method, a beads method, or a SDS method;   subjecting the genome DNA of the sample to random fragmenting by enzyme digestion, pulverization, ultrasound, or HydroShear method, to obtain DNA fragments;   and   subjecting the DNA fragments to single-end sequencing or pair-end sequencing, to obtain the reads of the DNA fragments.   
     
     
         10 . The method of  claim 1 , wherein the method further comprises:
 adding different indexes to each of the DNA fragments of the samples, to distinguish different samples.   
     
     
         11 . A system for detecting a copy number variation comprising:
 a reads obtaining unit, for obtaining reads from at least one part of a nucleic acid molecule of a sample;   a uniquely-mapped reads determining unit, for determining uniquely-mapped reads aligned to a genome reference sequence based on obtained reads;   a number of uniquely-mapped reads calculating unit, for dividing the genome reference sequence into a plurality of windows, and calculating the number of uniquely-mapped reads falling into each of the plurality of windows;   a number of uniquely-mapped reads correcting unit, for subjecting the number of uniquely-mapped reads falling into each of the plurality of windows to a GC correction, and to a correction based on the expected number of uniquely-mapped reads adjusted by a control set to obtain the corrected number of uniquely-mapped reads;   a candidate breakpoint selecting unit, for calculating a significance value of the difference between two numerical populations each consisting of the corrected numbers of uniquely-mapped reads falling into windows on each of the two sides of a demarcation point, the demarcation point being a starting point or an ending point of each of the plurality of windows, to thereby select the demarcation point having a smaller significance value as a candidate CNV breakpoint;   a breakpoint determining unit, for calculating a significance value of the difference between two numerical populations each consisting of the corrected number of uniquely-mapped reads falling into windows contained within each of two sequences, with one sequence ranging from a given candidate CNV breakpoint to an adjacent upstream candidate CNV breakpoint, and the other sequence ranging from the given candidate CNV breakpoint to an adjacent downstream candidate CNV breakpoint, and removing the candidate CNV breakpoint having the least significance at every turn and recalculating the significance value for the two candidate CNV breakpoints adjacent to the removed candidate CNV breakpoint, performing cyclic iteration until the significance values of all candidate CNV breakpoints are less than a termination threshold value, to thereby determine the CNV breakpoint.   
     
     
         12 . The system of  claim 11 , wherein each of the plurality of windows contains the same number of reference unique reads, or each of the plurality of windows has the same length. 
     
     
         13 . The system of  claim 11 , wherein the termination threshold value is obtained based on the control set consisting of normal samples. 
     
     
         14 . The system of  claim 11 , wherein the number of uniquely-mapped reads correcting unit comprises:
 a GC correction module, for grouping the plurality of windows based on a GC content, and obtaining a correction coefficient based on a mean value of the number of uniquely-mapped reads within one group and a mean value of the number of uniquely-mapped reads for all of the plurality of windows, and then subjecting the number of uniquely-mapped reads falling into each of the plurality of windows to the correction, to obtain the GC-corrected number of uniquely-mapped reads;   a window correction module, for calculating a ratio of the GC-corrected number of uniquely-mapped reads falling into each of the plurality of windows in the control set to the total number of uniquely-mapped reads; obtaining a mean value of the ratios for all windows corresponding to the control set; and calculating the expect number of uniquely-mapped reads for each of the plurality of windows in the sample based on the obtained mean value of the ratios and the total number of uniquely-mapped reads in the sample.   
     
     
         15 . The system of  claim 11 , wherein after the CNV breakpoint is determined by the breakpoint determining unit, the system further comprises:
 a breakpoint filtering unit, for determining a confidence interval of the normal corrected number of uniquely-mapped reads using the control set, based on a distribution of the corrected number of uniquely-mapped reads; and determining an abnormality presenting in the sequence between the two CNV breakpoints, if the mean value of the corrected number of uniquely-mapped reads within the sequence is outside the confidence interval.   
     
     
         16 . The system of  claim 15 , wherein the corrected number of uniquely-mapped reads fits a normal distribution, and the confidence interval is 95%. 
     
     
         17 . The system of  claim 11 , wherein the system further comprises one or more selected from the group consisting of:
 means for obtaining the sample from a human, with the sample comprising amniotic fluid obtained by amniocentesis, villus obtained by chorionic villi sampling, umbilical cord blood obtained by percutaneous umbilical blood sampling, spontaneous miscarrying fetus tissue or human peripheral blood;   means for obtaining a genomic DNA of the sample by a DNA extraction method such as a salting-out method, a column chromatography method, a beads method, or a SDS method;   means for subjecting the genome DNA of the sample to random fragmenting by enzyme digestion, pulverization, ultrasound, or HydroShear method, to obtain DNA fragments;   and   means for subjecting the DNA fragments to single-end sequencing or pair-end sequencing, to obtain the reads of the DNA fragments.   
     
     
         18 . The system of  claim 11 , wherein different samples are distinguished by adding different indexes to each of the DNA fragments of the samples. 
     
     
         19 . The system of  claim 11 , wherein in the candidate breakpoint selecting unit, a cyclization with a single chromosome or a whole genome is performed when selecting the candidate CNV breakpoint. 
     
     
         20 . The method of  claim 1 , wherein the sample is obtained from a subject in need of determining the copy number variation, and one or more of the steps of the method are computer-implemented.

Join the waitlist — get patent alerts

Track US2015056619A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.