US2021398616A1PendingUtilityA1
Methods and systems for aligning sequences in the presence of repeating elements
Assignee: SEVEN BRIDGES GENOMICS INCPriority: Oct 18, 2013Filed: Jun 25, 2021Published: Dec 23, 2021
Est. expiryOct 18, 2033(~7.2 yrs left)· nominal 20-yr term from priority
Inventors:Deniz Kural
G16B 30/10C12Q 1/6869G16B 30/00C12Q 2537/165C12Q 2535/122
71
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The invention includes methods for aligning reads (e.g., nucleic acid reads) comprising repeating sequences, methods for building reference sequence constructs comprising repeating sequences, and systems that can be used to align reads comprising repeating sequences. The method is scalable, and can be used to align millions of reads to a construct thousands of bases long. The methods and systems can additionally account for variability within a repeating sequence, or near to a repeating sequence, due to genetic mutation.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method, comprising:
using at least one computer hardware processor to perform:
obtaining a first sequence read and a second sequence read, wherein the second sequence read comprises at least a portion of a repetitive sequence within a genome and is separated from the first sequence read by a first distance;
obtaining a reference sequence construct represented as a directed acyclic graph (DAG) data structure, the reference sequence construct comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored as a respective string of one or more symbols, wherein the DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols;
aligning the first sequence read to the reference sequence construct at least in part by:
determining a first plurality of scores corresponding to a respective first plurality of alignments between the first sequence read and the reference sequence construct, the first plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the reference sequence construct, the first score being determined based on a degree of overlap between the first sequence read and the first string and the first sequence read and the second string; and
aligning the second sequence read to the reference sequence construct based on the first distance and a result of aligning the first sequence read to the reference sequence construct.
22 . The method of claim 21 , further comprising identifying, based on a result of aligning the second sequence read to the reference sequence construct, the repetitive sequence as present in the genome.
23 . The method of claim 21 , wherein aligning the second sequence read to the reference sequence construct comprises:
determining a second plurality of scores corresponding to a respective plurality of alignments between the second sequence read and the reference sequence construct, the second plurality of scores including a second score corresponding to a second alignment between the second sequence read and at least a second portion of the reference sequence construct, the second score being determined based on a degree of overlap between the second sequence read and the first string and the second sequence read and the second string.
24 . The method of claim 21 ,
wherein aligning the first sequence read comprises determining an aligned position of the first sequence read with respect to the reference sequence construct, and wherein aligning the second sequence read comprises:
determining a position of the second sequence read with respect to the reference sequence construct; and
determining a distance between the aligned position of the first sequence read and the position of the second sequence read.
25 . The method of claim 21 ,
wherein aligning the first sequence read comprises determining an aligned position of the first sequence read with respect to the reference sequence construct, wherein aligning the second sequence read comprises:
determining a position of the second sequence read with respect to the reference sequence construct;
determining whether the aligned position of the first sequence read and the position of the second sequence read are separated by the first distance; and
excluding the position of the second sequence read as an aligned position of the second sequence read when the position of the second sequence read and the aligned position of the first sequence read are not separated by the first distance.
26 . The method of claim 21 , wherein the first sequence read and the second sequence read are paired mates.
27 . The method of claim 21 , further comprising obtaining a third sequence read, wherein the second sequence read and the third sequence read are separated by a second distance, and wherein aligning the second sequence read to the reference sequence construct comprises aligning the second sequence read based on the second distance.
28 . The method of claim 27 , wherein the second sequence read and the third sequence read are paired mates.
29 . The method of claim 21 , further comprising identifying the repetitive sequence as present at an aligned position of the second sequence read, wherein the aligned position of the second sequence read is a result of aligning the second sequence read to the reference sequence construct.
30 . The method of claim 21 , wherein at least two positions in the reference sequence construct represent the repetitive sequence.
31 . The method of claim 21 , wherein the first sequence read and the second sequence read were previously obtained from a genetic sample, the method further comprising:
determining a distribution of repetitive sequences in the genetic sample based on results of aligning the first sequence read and the second sequence read to the reference sequence construct.
32 . The method of claim 21 , wherein the first sequence read and the second sequence read were previously obtained from a genetic sample, the method further comprising:
determining a genotype for the genetic sample based on results of aligning the first sequence read and the second sequence read to the reference sequence construct.
33 . A system, comprising:
at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining a first sequence read and a second sequence read, wherein the second sequence read comprises at least a portion of a repetitive sequence within a genome and is separated from the first sequence read by a first distance;
obtaining a reference sequence construct represented as a directed acyclic graph (DAG) data structure, the reference sequence construct comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored as a respective string of one or more symbols, wherein the DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols;
aligning the first sequence read to the reference sequence construct at least in part by:
determining a first plurality of scores corresponding to a respective first plurality of alignments between the first sequence read and the reference sequence construct, the first plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the reference sequence construct, the first score being determined based on a degree of overlap between the first sequence read and the first string and the first sequence read and the second string; and
aligning the second sequence read to the reference sequence construct based on the first distance and a result of aligning the first sequence read to the reference sequence construct.
34 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining a first sequence read and a second sequence read, wherein the second sequence read comprises at least a portion of a repetitive sequence within a genome and is separated from the first sequence read by a first distance; obtaining a reference sequence construct represented as a directed acyclic graph (DAG) data, the reference sequence construct comprising a plurality of nodes, wherein a node of the plurality of nodes represents a nucleotide sequence of one or more nucleotides stored as a respective string of one or more symbols, wherein the DAG data structure further comprises a first node connected by edges to a first alternate node, the first node representing a first nucleotide sequence stored as a first string of one or more symbols and the first alternate node representing a second nucleotide sequence stored as a second string of one or more symbols;
aligning the first sequence read to the reference sequence construct at least in part by:
determining a first plurality of scores corresponding to a respective first plurality of alignments between the first sequence read and the reference sequence construct, the first plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the reference sequence construct, the first score being determined based on a degree of overlap between the first sequence read and the first string and the first sequence read and the second string; and
aligning the second sequence read to the reference sequence construct based on the first distance and a result of aligning the first sequence read to the reference sequence construct.Join the waitlist — get patent alerts
Track US2021398616A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.