US2022411881A1PendingUtilityA1
Methods and systems for identifying disease-induced mutations
Assignee: SEVEN BRIDGES GENOMICS INCPriority: Oct 18, 2013Filed: Jun 29, 2022Published: Dec 29, 2022
Est. expiryOct 18, 2033(~7.2 yrs left)· nominal 20-yr term from priority
Inventors:Deniz Kural
C12Q 1/6886C12Q 1/6883G16B 20/00C12Q 2600/112G16B 30/10G16B 20/20C12Q 2600/156C12Q 2600/16
75
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The invention includes methods and systems for identifying diseased-induced mutations by producing multi-dimensional reference sequence constructs that account for variations between individuals, different diseases, and different stages of those diseases. Once constructed, these reference sequence constructs can be used to align sequence reads corresponding to genetic samples from patients suspected of having a disease, or who have had the disease and are in suspected remission. The reference sequence constructs also provide insight to the genetic progression of the disease.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method for identifying one or more mutations associated with an advanced stage of cancer, the method comprising:
using at least one computer hardware processor to perform:
obtaining, from at least one non-transitory computer-readable storage medium, a graph data structure representing a reference sequence and genetic variation of the reference sequence, wherein a first path through the graph data structure represents a first sequence associated with a first string of one or more symbols and a second path through the graph data structure represents a second sequence associated with a second string of one or more symbols, the first sequence representing at least a first portion of the reference sequence and the second sequence representing a first genetic variation from the at least the first portion of the reference sequence, the first genetic variation being associated with the cancer;
obtaining a plurality of sequence reads previously obtained from a genetic sample from a subject, wherein the genetic sample is associated with the advanced stage of the cancer;
aligning a first sequence read of the plurality of sequence reads to the graph data structure, the aligning comprising:
determining a plurality of scores corresponding to a respective plurality of alignments between the first sequence read and the graph data structure, the plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the graph data structure, the first score being determined based on a degree of overlap between the first sequence read and the first string and a degree of overlap between the first sequence read and the second string; and
identifying one or more differences between the first sequence read and at least a portion of the graph data structure against which the first sequence read is aligned, wherein the one or more differences represent the one or more mutations associated with the advanced stage of the cancer.
22 . The method of claim 21 , further comprising:
updating the graph data structure, the updating comprising:
representing, in the graph data structure, the identified one or more differences between the first sequence read and the at least the portion of the graph data structure.
23 . The method of claim 22 , wherein a difference of the one or more differences comprises a third sequence associated with a third string of one or more symbols, and wherein representing the difference in the graph data structure comprises representing the difference as a third path through the graph data structure.
24 . The method of claim 22 , further comprising:
obtaining a second plurality of sequence reads previously obtained from a second genetic sample; and aligning at least one sequence read of the second plurality of sequence reads to the updated graph data structure.
25 . The method of claim 21 , wherein the at least the portion of the graph data structure represents a sequence associated with a string of one or more symbols, and
where identifying the one or more differences between the first sequence read and the at least the portion of the graph data structure comprises:
comparing the first sequence read to the sequence represented by the at least the portion of the graph data structure.
26 . The method of claim 21 , wherein identifying the one or more differences between the first sequence read and the at least the portion of the graph reference comprises:
determining whether the at least the portion of the graph data structure comprises at least a portion of the second path through the graph data structure; and determining that the one or more differences represent the one or more mutations associated with the advanced stage of the cancer when the at least the portion of the graph data structure comprises the at least the portion of the second path.
27 . The method of claim 21 , wherein the genetic sample comprises a tumor sample.
28 . The method of claim 21 , wherein the second sequence is associated with a second genetic sample from the subject, the second genetic sample being associated with the cancer.
29 . The method of claim 28 , wherein a third path through the graph data structure represents a third sequence, the third sequence representing a second genetic variation from the at least the first portion of the reference sequence, wherein the second genetic variation represents a non-cancerous genetic variation.
30 . The method of claim 29 , wherein the third sequence is associated with a third genetic sample from the subject, the third genetic sample being non-cancerous.
31 . The method of claim 21 , further comprising:
aligning a second sequence read of the plurality of sequence reads to the graph data structure; and identifying one or more second differences between the second sequence read and at least a second portion of the graph data structure against which the second sequence read is aligned, wherein the one or more second differences represent one or more second mutations associated with the advanced stage of the cancer.
32 . The method of claim 21 , further comprising:
determining, based on a presence of the one or more mutations associated with the advanced stage of the cancer, whether the subject should undergo additional testing.
33 . The method of claim 21 , wherein the genetic variation comprises an insertion, a deletion, a polymorphism, or a structural variant.
34 . The method of claim 21 , wherein the graph data structure comprises a directed acyclic graph (DAG) data structure.
35 . The method of claim 21 , wherein the graph data structure comprises a plurality of nodes, and wherein the first path comprises a node of the plurality of nodes not included in the second path.
36 . The method of claim 21 , wherein the graph data structure comprises a plurality of nodes, and wherein the second path comprises a node of the plurality of nodes not included in the first path.
37 . The method of claim 21 , wherein the reference sequence comprises a reference genome, and wherein the genetic variation comprises a genetic variation of the reference genome.
38 . The method of claim 21 , further comprising providing information indicative of the one or more differences to a health care provider.
39 . A system comprising:
at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform:
obtaining, from at least one non-transitory computer-readable storage medium, a graph data structure representing a reference sequence and genetic variation of the reference sequence, wherein a first path through the graph data structure represents a first sequence associated with a first string of one or more symbols and a second path through the graph data structure represents a second sequence associated with a second string of one or more symbols, the first sequence representing at least a first portion of the reference sequence and the second sequence representing a first genetic variation from the at least the first portion of the reference sequence, the first genetic variation being associated with the cancer;
obtaining a plurality of sequence reads previously-obtained from a genetic sample, wherein the genetic sample is associated with the advanced stage of the cancer;
aligning a first sequence read of the plurality of sequence reads to the graph data structure, the aligning comprising:
determining a plurality of scores corresponding to a respective plurality of alignments between the first sequence read and the graph data structure, the plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the graph data structure, the first score being determined based on a degree of overlap between the first sequence read and the first string and a degree of overlap between the first sequence read and the second string; and
identifying, based on results of the aligning, one or more differences between the first sequence read and at least a portion of the graph data structure, wherein the one or more differences represent the one or more mutations associated with the advanced stage of the cancer.
40 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer-hardware processor to perform:
obtaining, from at least one non-transitory computer-readable storage medium, a graph data structure representing a reference sequence and genetic variation of the reference sequence, wherein a first path through the graph data structure represents a first sequence associated with a first string of one or more symbols and a second path through the graph data structure represents a second sequence associated with a second string of one or more symbols, the first sequence representing at least a first portion of the reference sequence and the second sequence representing a first genetic variation from the at least the first portion of the reference sequence, the first genetic variation being associated with the cancer; obtaining a plurality of sequence reads previously-obtained from a genetic sample, wherein the genetic sample is associated with the advanced stage of the cancer;
aligning a first sequence read of the plurality of sequence reads to the graph data structure, the aligning comprising:
determining a plurality of scores corresponding to a respective plurality of alignments between the first sequence read and the graph data structure, the plurality of scores including a first score corresponding to a first alignment between the first sequence read and at least a first portion of the graph data structure, the first score being determined based on a degree of overlap between the first sequence read and the first string and a degree of overlap between the first sequence read and the second string; and identifying, based on results of the aligning, one or more differences between the first sequence read and at least a portion of the graph data structure, wherein the one or more differences represent the one or more mutations associated with the advanced stage of the cancer.Join the waitlist — get patent alerts
Track US2022411881A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.