US2021280272A1PendingUtilityA1

Methods and systems for quantifying sequence alignment

Assignee: SEVEN BRIDGES GENOMICS INCPriority: Oct 18, 2013Filed: Nov 2, 2020Published: Sep 9, 2021
Est. expiryOct 18, 2033(~7.2 yrs left)· nominal 20-yr term from priority
Inventors:Deniz Kural
G16B 30/10G16B 30/00G16B 50/00G16B 45/00C12Q 2537/165G16B 30/20G16B 20/00G16B 20/20
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The invention includes methods for aligning reads (e.g., nucleic acid reads, amino acid reads) to a reference sequence construct, methods for building the reference sequence construct, and systems that use the alignment methods and constructs to produce sequences. The invention also includes methods and systems for evaluating the quality of the alignment between the reads and the reference sequence construct. The method is scalable, and can be used to align millions of reads to a construct thousands of bases or amino acids long. The invention additionally includes methods for identifying a disease or a genotype based upon alignment of nucleic acid reads to a location in the construct.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A system, comprising:
 at least one computer hardware processor; and   at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:
 obtaining a sequence read that has been previously obtained from a genetic sample; 
 obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation; 
 aligning the sequence read to the reference DAG data structure at least in part by:
 determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string; 
 
 determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and 
 genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold. 
   
     
     
         22 . The system of  claim 21 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
 wherein determining the first score further comprises determining overlaps between the sequence read and the third string.   
     
     
         23 . The system of  claim 22 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string. 
     
     
         24 . The system of  claim 21 , wherein determining the overlap value comprises:
 determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.   
     
     
         25 . The system of  claim 24 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols. 
     
     
         26 . The system of  claim 21 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation. 
     
     
         27 . The system of  claim 21 , wherein aligning the sequence read to the reference DAG data structure further comprises:
 creating, in the at least one non-transitory computer-readable storage medium, a first matrix for the first node and a second matrix for the first alternate node, the first matrix representing one or more alignments between the sequence read and the first string and the second matrix representing one or more alignments between the sequence read and the second string; and   wherein determining the plurality of scores comprises determining a first plurality of scores for the first matrix and a second plurality of scores for the second matrix, the first plurality of scores corresponding to the one or more alignments between the sequence read and the first string and the second plurality of scores corresponding to the one or more alignments between the sequence read and the second string.   
     
     
         28 . A method, comprising:
 obtaining a sequence read that has been previously obtained from a genetic sample;   obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation;   aligning the sequence read to the reference DAG data structure at least in part by:
 determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string; 
   determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and   genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold.   
     
     
         29 . The method of  claim 28 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
 wherein determining the first score further comprises determining overlaps between the sequence read and the third string.   
     
     
         30 . The method of  claim 29 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string. 
     
     
         31 . The method of  claim 28 , wherein determining the overlap value comprises:
 determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.   
     
     
         32 . The method of  claim 31 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols. 
     
     
         33 . The method of  claim 28 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation. 
     
     
         34 . The method of  claim 28 , wherein aligning the sequence read to the reference DAG data structure further comprises:
 creating, in the at least one non-transitory computer-readable storage medium, a first matrix for the first node and a second matrix for the first alternate node, the first matrix representing one or more alignments between the sequence read and the first string and the second matrix representing one or more alignments between the sequence read and the second string; and   wherein determining the plurality of scores comprises determining a first plurality of scores for the first matrix and a second plurality of scores for the second matrix, the first plurality of scores corresponding to the one or more alignments between the sequence read and the first string and the second plurality of scores corresponding to the one or more alignments between the sequence read and the second string.   
     
     
         35 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform:
 obtaining a sequence read that has been previously obtained from a genetic sample;   obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation;   aligning the sequence read to the reference DAG data structure at least in part by:
 determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string; 
   determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and   genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold.   
     
     
         36 . The at least one non-transitory computer-readable storage medium of  claim 35 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
 wherein determining the first score further comprises determining overlaps between the sequence read and the third string.   
     
     
         37 . The at least one non-transitory computer-readable storage medium of  claim 36 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string. 
     
     
         38 . The at least one non-transitory computer-readable storage medium of  claim 35 , wherein determining the overlap value comprises:
 determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.   
     
     
         39 . The at least one non-transitory computer-readable storage medium of  claim 38 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols. 
     
     
         40 . The at least one non-transitory computer-readable storage medium of  claim 35 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation.

Join the waitlist — get patent alerts

Track US2021280272A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.