Somatic variant cooccurrence with abnormally methylated fragments
Abstract
Systems and methods for identifying variant alleles as somatic or germline are provided. Reference and variant alleles for a genomic position are identified. Methylation states and sequences of nucleic acid fragment sequences that map to the genomic position are obtained from a sample of a subject. Using the sequences of nucleic acid fragment sequences, each nucleic acid fragment sequence that has the reference allele is assigned to a reference subset, and each nucleic acid fragment sequence that has the variant allele is assigned to a variant subset. One or more indications of the methylation states across the nucleic acid fragment sequences in the variant subset and an indication of the number of nucleic acid fragment sequences in the reference subset versus the variant subset are applied to a trained binary classifier. An identification of the variant allele at the genomic position as somatic or germline is obtained from the classifier.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . A method of identifying a variant allele at a genomic position in a test subject as somatic or germline, the method comprising:
obtaining an identification of a reference allele at the genomic position; obtaining an identification of the variant allele at the genomic position; obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in a respective plurality of nucleic acid fragment sequences in a sequencing dataset derived from a liquid biological sample obtained from the test subject that map onto the genomic position, wherein the sequencing dataset comprises at least 1×10 6 nucleic acid fragment sequences; using (i) the identification of the reference allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the reference allele, at the genomic position, to a reference subset; using (i) the identification of the variant allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the variant allele, at the genomic position, to a variant subset; and applying, to a trained binary classifier, at least (i) one or more indications of methylation state across the methylation state of each nucleic acid fragment sequence in the variant subset and (ii) an indication of a number of nucleic acid fragment sequences in the reference subset versus a number of nucleic acid fragment sequences in the variant subset, wherein the trained binary classifier comprises at least 10 parameters, thereby obtaining from the trained binary classifier an identification of the variant allele at the genomic position in the test subject as somatic or germline.
2 . The method of claim 1 , wherein the method further comprises:
inputting a reference genome into a computer system comprising a processor coupled to a non-transitory memory, and using the computer system to determine that each respective nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences maps to the genomic position by aligning the respective nucleic acid fragment sequence to the reference genome.
3 . The method of claim 1 , wherein
a first nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences has a plurality of CpG sites; wherein the first nucleic acid fragment sequence has a corresponding methylation pattern across the plurality of CpG sites; wherein the methylation state of the first nucleic acid fragment sequence is a p-value, and wherein the method further comprises:
determining the p-value of the first nucleic acid fragment sequence, at least in part, by comparison of the corresponding methylation pattern of the first nucleic acid fragment sequence to a corresponding distribution of methylation patterns of those nucleic acid fragment sequences in a healthy noncancer cohort dataset that each have the respective plurality of CpG sites.
4 . The method of claim 1 , wherein the variant allele is an insertion, a deletion, or a single nucleotide polymorphism.
5 . The method of claim 1 , wherein, when the variant allele at the genomic position is determined by the trained binary classifier to be germline, the method further comprises:
using the variant allele in the test subject to perform an action selected from the group consisting of: determining a cancer risk of the test subject, predicting an ethnicity of the test subject, and determining a tumor fraction of the test subject.
6 . The method of claim 1 , wherein:
each indication in the one or more indications of methylation state across the variant subset is: a measure of central tendency of a methylation state p-value across the variant subset, a minimum methylation state p-value across the variant subset, a maximum methylation state p-value across the variant subset, or a measure of spread of a methylation state p-value across the variant subset.
7 . The method of claim 1 , wherein the one or more indications of methylation state across the variant subset is a plurality of indications of methylation state across the variant subset comprising at least 2, at least 3, or all four of:
a measure of central tendency of a methylation state p-value across the variant subset, a minimum methylation state p-value across the variant subset, a maximum methylation state p-value across the variant subset, and a measure of spread of a methylation state p-value across the variant subset.
8 . The method of claim 1 , wherein the applying, to the trained binary classifier, further applies one of:
one or more CpG site indications across the variant subset; one or more indications of methylation state across the reference subset; or one or more CpG site indications across the reference subset.
9 . The method of claim 1 , wherein the obtaining an identification of the variant allele at the genomic position comprises determining that the respective plurality of nucleic acid fragments support a variant allele call at the genomic position.
10 . The method of claim 1 , further comprising performing methylation sequencing to obtain the methylation state and the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences.
11 . A computing system, comprising:
one or more processors; memory storing one or more programs to be executed by the one or more processor, the one or more programs comprising instructions for calling a variant at a genomic position in a test subject by a method comprising: obtaining an identification of a reference allele at the genomic position; obtaining an identification of the variant allele at the genomic position; obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in a respective plurality of nucleic acid fragment sequences in a sequencing dataset derived from a liquid biological sample obtained from the test subject that map onto the genomic position, wherein the sequencing dataset comprises at least 10{circumflex over ( )}6 nucleic acid fragment sequences; using (i) the identification of the reference allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the reference allele, at the genomic position, to a reference subset; using (i) the identification of the variant allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the variant allele, at the genomic position, to a variant subset; and applying, to a trained binary classifier, at least (i) one or more indications of methylation state across the methylation state of each nucleic acid fragment sequence in the variant subset and (ii) an indication of a number of nucleic acid fragment sequences in the reference subset versus a number of nucleic acid fragment sequences in the variant subset, wherein the trained binary classifier comprises at least 10 parameters, thereby obtaining from the trained binary classifier an identification of the variant allele at the genomic position in the test subject as somatic or germline.
12 . The computing system of claim 11 , wherein the instructions further comprise:
inputting a reference genome into a computer system comprising a processor coupled to a non-transitory memory, and using the computer system to determine that each respective nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences maps to the genomic position by aligning the respective nucleic acid fragment sequence to the reference genome.
13 . The computing system of claim 11 , wherein:
a first nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences has a plurality of CpG sites; wherein the first nucleic acid fragment sequence has a corresponding methylation pattern across the plurality of CpG sites; wherein the methylation state of the first nucleic acid fragment sequence is a p-value, and wherein the instructions further comprise:
determining the p-value of the first nucleic acid fragment sequence, at least in part, by comparison of the corresponding methylation pattern of the first nucleic acid fragment sequence to a corresponding distribution of methylation patterns of those nucleic acid fragment sequences in a healthy noncancer cohort dataset that each have the respective plurality of CpG sites.
14 . The computing system of claim 11 , wherein, when the variant allele at the genomic position is determined by the trained binary classifier to be germline, the instructions further comprise:
using the variant allele in the test subject to perform an action selected from the group consisting of: determining a cancer risk of the test subject, predicting an ethnicity of the test subject, and determining a tumor fraction of the test subject.
15 . The computing system of claim 11 , wherein:
each indication in the one or more indications of methylation state across the variant subset is: a measure of central tendency of a methylation state p-value across the variant subset, a minimum methylation state p-value across the variant subset, a maximum methylation state p-value across the variant subset, or a measure of spread of a methylation state p-value across the variant subset.
16 . The computing system of claim 11 , wherein the one or more indications of methylation state across the variant subset is a plurality of indications of methylation state across the variant subset comprising at least 2, at least 3, or all four of:
a measure of central tendency of a methylation state p-value across the variant subset, a minimum methylation state p-value across the variant subset, a maximum methylation state p-value across the variant subset, and a measure of spread of a methylation state p-value across the variant subset.
17 . The computing system of claim 11 , wherein the instruction to apply, to the trained binary classifier, further comprise instructions to apply one of:
one or more CpG site indications across the variant subset; one or more indications of methylation state across the reference subset; or one or more CpG site indications across the reference subset.
18 . The computing system of claim 11 , wherein the instructions to obtain an identification of the variant allele at the genomic position comprise instructions to determine that the respective plurality of nucleic acid fragments support a variant allele call at the genomic position.
19 . The computing system of claim 11 , wherein the instructions further comprise performing methylation sequencing to obtain the methylation state and the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences.
20 . A non-transitory computer readable storage medium storing one or more programs for calling a variant at a genomic position in a test subject, the one or more programs configured for execution by a computer, wherein the one or more programs comprise instructions for:
obtaining an identification of a reference allele at the genomic position; obtaining an identification of the variant allele at the genomic position; obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in a respective plurality of nucleic acid fragment sequences in a sequencing dataset derived from a liquid biological sample obtained from the test subject that map onto the genomic position, wherein the sequencing dataset comprises at least 10{circumflex over ( )}6 nucleic acid fragment sequences; using (i) the identification of the reference allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the reference allele, at the genomic position, to a reference subset; using (i) the identification of the variant allele at the genomic position and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the variant allele, at the genomic position, to a variant subset; and applying, to a trained binary classifier, at least (i) one or more indications of methylation state across the methylation state of each nucleic acid fragment sequence in the variant subset and (ii) an indication of a number of nucleic acid fragment sequences in the reference subset versus a number of nucleic acid fragment sequences in the variant subset, wherein the trained binary classifier comprises at least 10 parameters, thereby obtaining from the trained binary classifier an identification of the variant allele at the genomic position in the test subject as somatic or germline.
21 . A method of training a classifier to identify a variant allele at a genomic position in a test subject as somatic or germline, the method comprising:
A) obtaining an identification of a reference allele at the genomic position; B) for each respective subject in a plurality of subjects, for each respective genomic position in a plurality of genomic positions, performing a procedure comprising:
i) obtaining an orthogonal call for the variant allele at the respective genomic position as one of somatic or germline for the respective subject;
ii) obtaining an identification of the variant allele at the respective genomic position for the respective subject;
iii) obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in a respective plurality of nucleic acid fragment sequences in a sequencing dataset derived from a liquid biological sample obtained from the respective subject that map onto the respective genomic position, wherein the sequencing dataset comprises at least 1×10 6 nucleic acid fragment sequences;
iv) using (a) the identification of the reference allele at the respective genomic position and (b) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the reference allele, at the respective genomic position, to a reference subset;
v) using (a) the identification of the variant allele at the respective genomic position and (b) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences to assign each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences that has the variant allele, at the respective genomic position, to a variant subset; and
C) using, for each respective subject in the plurality of subjects, for each respective genomic position in the plurality of genomic positions, at least (i) one or more indications of methylation state across the methylation state of each nucleic acid fragment sequence in the variant subset for the respective subject for the respective genomic position (ii) an indication of a number of nucleic acid fragment sequences in the reference subset versus a number of nucleic acid fragment sequences in the variant subset for the respective subject for the respective genomic position and (iii) the orthogonal call for the variant allele at the respective genomic position as one of somatic or germline for the respective subject to train the classifier to identify a variant allele at a genomic position in a test subject as somatic or germline, wherein the classifier comprises at least 10 parameters.Join the waitlist — get patent alerts
Track US2023057154A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.