Methods and systems for quantifying sequence alignment
Abstract
The invention includes methods for aligning reads (e.g., nucleic acid reads, amino acid reads) to a reference sequence construct, methods for building the reference sequence construct, and systems that use the alignment methods and constructs to produce sequences. The invention also includes methods and systems for evaluating the quality of the alignment between the reads and the reference sequence construct. The method is scalable, and can be used to align millions of reads to a construct thousands of bases or amino acids long. The invention additionally includes methods for identifying a disease or a genotype based upon alignment of nucleic acid reads to a location in the construct.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system, comprising:
at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining a sequence read that has been previously obtained from a genetic sample;
obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation;
aligning the sequence read to the reference DAG data structure at least in part by:
determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string;
determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and
genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold.
22 . The system of claim 21 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
wherein determining the first score further comprises determining overlaps between the sequence read and the third string.
23 . The system of claim 22 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string.
24 . The system of claim 21 , wherein determining the overlap value comprises:
determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.
25 . The system of claim 24 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols.
26 . The system of claim 21 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation.
27 . The system of claim 21 , wherein aligning the sequence read to the reference DAG data structure further comprises:
creating, in the at least one non-transitory computer-readable storage medium, a first matrix for the first node and a second matrix for the first alternate node, the first matrix representing one or more alignments between the sequence read and the first string and the second matrix representing one or more alignments between the sequence read and the second string; and wherein determining the plurality of scores comprises determining a first plurality of scores for the first matrix and a second plurality of scores for the second matrix, the first plurality of scores corresponding to the one or more alignments between the sequence read and the first string and the second plurality of scores corresponding to the one or more alignments between the sequence read and the second string.
28 . A method, comprising:
obtaining a sequence read that has been previously obtained from a genetic sample; obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation; aligning the sequence read to the reference DAG data structure at least in part by:
determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string;
determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold.
29 . The method of claim 28 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
wherein determining the first score further comprises determining overlaps between the sequence read and the third string.
30 . The method of claim 29 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string.
31 . The method of claim 28 , wherein determining the overlap value comprises:
determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.
32 . The method of claim 31 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols.
33 . The method of claim 28 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation.
34 . The method of claim 28 , wherein aligning the sequence read to the reference DAG data structure further comprises:
creating, in the at least one non-transitory computer-readable storage medium, a first matrix for the first node and a second matrix for the first alternate node, the first matrix representing one or more alignments between the sequence read and the first string and the second matrix representing one or more alignments between the sequence read and the second string; and wherein determining the plurality of scores comprises determining a first plurality of scores for the first matrix and a second plurality of scores for the second matrix, the first plurality of scores corresponding to the one or more alignments between the sequence read and the first string and the second plurality of scores corresponding to the one or more alignments between the sequence read and the second string.
35 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining a sequence read that has been previously obtained from a genetic sample; obtaining, from the at least one non-transitory computer-readable storage medium, a reference directed acyclic graph (DAG) data structure comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored in the at least one non-transitory computer-readable storage medium as a respective string of one or more symbols, wherein the reference DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols representing a genetic structural variation; aligning the sequence read to the reference DAG data structure at least in part by:
determining a plurality of scores corresponding to a respective plurality of alignments between the sequence read and the reference DAG data structure, the plurality of scores including a first score corresponding to a first alignment between the sequence read and at least a portion of the reference DAG data structure, the first score being determined based on a degree of overlap between the sequence read and the first string and a degree of overlap between the sequence read and the second string;
determining, based on results of the aligning, an overlap value indicative of a number of overlapping symbols between the sequence read and one of the first string or the second string; and genotyping the genetic sample based on the aligning when the overlap value exceeds a threshold.
36 . The at least one non-transitory computer-readable storage medium of claim 35 , wherein the DAG data structure further comprises a second alternate node representing a third nucleotide sequence stored as a third string of one or more symbols; and
wherein determining the first score further comprises determining overlaps between the sequence read and the third string.
37 . The at least one non-transitory computer-readable storage medium of claim 36 , wherein the overlap value is indicative of the number of overlapping symbols between the sequence read and one of the first string, the second string, or the third string.
38 . The at least one non-transitory computer-readable storage medium of claim 35 , wherein determining the overlap value comprises:
determining a first number of overlapping symbols between the sequence read and the first string and a second number of overlapping symbols between the sequence read and the second string.
39 . The at least one non-transitory computer-readable storage medium of claim 38 , wherein determining the overlap value further comprises identifying a smallest number of overlapping symbols from among the first number of overlapping symbols and the second number of overlapping symbols.
40 . The at least one non-transitory computer-readable storage medium of claim 35 , wherein genotyping the genetic sample based on the aligning comprises genotyping the genetic sample with respect to the genetic structural variation.Join the waitlist — get patent alerts
Track US2021280272A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.