US2022383977A1PendingUtilityA1

Methods and processes for non-invasive assessment of genetic variations

Assignee: SEQUENOM INCPriority: Oct 6, 2011Filed: Aug 2, 2022Published: Dec 1, 2022
Est. expiryOct 6, 2031(~5.2 yrs left)· nominal 20-yr term from priority
G16B 20/10G16B 30/20C12Q 2600/156G16B 20/20G16B 20/00G16B 30/10G16B 20/40G16B 30/00C12Q 1/6883
79
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Provided herein are methods, processes and apparatuses for non-invasive assessment of genetic variations.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising one or more processors and memory,
 which memory comprises instructions executable by the one or more processors and which memory comprises counts of sequence reads mapped to portions of a reference genome, which sequence reads are reads of circulating cell-free nucleic acid from a test sample from a pregnant female; and   which instructions executable by the one or more processors are configured to:   (a) reduce error in the counts of the sequence reads, wherein the error is reduced according to a process comprising:
 (1) assigning a guanine and cytosine (GC) bias coefficient to the test sample based on a first fitted relation between (i) the counts of the sequence reads mapped to each of the portions and (ii) GC content for each of the portions, wherein the GC bias coefficient is a slope for a linear fitted relation for the first fitted relation or a curvature estimation for a non-linear fitted relation for the first fitted relation; and 
 (2) calculating a genomic section level for each of the portions for the test sample based on the counts of the sequence reads, the GC bias coefficient of (a)(1), and a second fitted relation, for each of the portions, between (i) the GC bias coefficient for each of multiple samples and (ii) the counts of the sequence reads mapped to each of the portions for the multiple samples, thereby providing calculated genomic section levels, whereby error in the counts of the sequence reads is reduced; and 
   (b) output a classification of the presence or absence of a copy number variation for the test sample according to the calculated genomic section levels, wherein a measure of deviation between (i) calculated genomic section levels for portions that include a copy number variation and (ii) expected genomic section levels for portions that include no copy number variation is larger for counts with error reduction according to (a) than a measure of deviation between (b)(i) and (b)(ii) for counts with no error reduction according to (a).   
     
     
         2 . The system of  claim 1 , wherein the GC bias coefficient in (a)(1) is a slope for a linear fitted relation for the first fitted relation determined by linear regression. 
     
     
         3 . The system of  claim 1 , wherein the GC bias coefficient in (a)(1) is a curvature estimation determined by a non-linear fitted relation for the first fitted relation. 
     
     
         4 . The system of  claim 1 , wherein the second fitted relation of (a)(2) is linear. 
     
     
         5 . The system of  claim 4 , wherein a slope of the second fitted relation in (a)(2) is determined by linear regression. 
     
     
         6 . The system of  claim 5 , wherein the GC bias coefficient for each of the multiple samples in (a)(2)( i ) is the slope of a third fitted linear relation, for each of the multiple samples, between (i) the counts of the sequence reads mapped to each of the portions and (ii) GC content for each of the portions. 
     
     
         7 . The system of  claim 6 , wherein a calculated genomic section level L is calculated for the test sample for each portion of the reference genome according to Equation B:
     L =( M−GS )/ I   Equation B
   
       wherein M is the counts of the sequence reads mapped to the portion for the test sample, G is the GC bias coefficient for the test sample, I is an intercept of the second fitted linear relation of (a)(2) for the portion, and S is a slope of the second fitted linear relationship of (a)(2) for the portion. 
     
     
         8 . The system of  claim 1 , wherein the instructions executable by the one or more processors are further configured to generate at least one Z-score from the calculated genomic section levels. 
     
     
         9 . The system of  claim 1 , wherein the instructions executable by the one or more processors are further configured to filter one or more portions and remove counts associated with the one or more portions for the classification of the presence or absence of the copy number variation in (b). 
     
     
         10 . The system of  claim 9 , wherein the instructions executable by the one or more processors are configured to (i) normalize the counts of sequence reads according to (a), thereby generating normalized counts, and removing normalized counts associated with one or more filtered portions, thereby yielding filtered normalized counts; or (ii) remove counts of sequence reads associated with one or more filtered portions prior to (a), and normalize the counts in portions that were not removed according to (a), thereby yielding filtered normalized counts. 
     
     
         11 . The system of  claim 9 , wherein the one or more filtered portions are selected according to one or more criteria chosen from portions having no guanosine and cytosine (GC) content, portions consistently receiving no counts, and repeat masking. 
     
     
         12 . The system of  claim 9 , wherein the one or more filtered portions are selected according to one or more criteria chosen from measure of error or mappability, or measure of error and mappability. 
     
     
         13 . The system of  claim 12 , wherein the measure of error is an R factor. 
     
     
         14 . The system of  claim 13 , wherein portions of the reference genome having an R factor of about 7% or greater were selected as filtered portions. 
     
     
         15 . The system of  claim 13 , wherein portions of the reference genome having an R factor of about 7% to about 10% were selected as filtered portions. 
     
     
         16 . The system of  claim 1 , wherein the sequence reads for the test sample were generated by a genome-wide massively parallel sequencing process. 
     
     
         17 . The system of  claim 16 , wherein the sequencing is at about 1-fold coverage or less. 
     
     
         18 . The system of  claim 16 , wherein the sequencing is at about 1-fold coverage or greater. 
     
     
         19 . The system of  claim 16 , wherein the instructions executable by the one or more processors are further configured to map the sequence reads to the portions of the reference genome, and count the mapped sequence reads, thereby generating the counts of the sequence reads mapped to the portions of the reference genome. 
     
     
         20 . The system of  claim 1 , wherein the nucleic acid is from blood plasma or blood serum. 
     
     
         21 . The system of  claim 1 , wherein the copy number variation is chosen from a chromosome aneuploidy, a microdeletion, and a microduplication.

Join the waitlist — get patent alerts

Track US2022383977A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.