Determination of base modifications of nucleic acids
Abstract
Systems and methods for using determination of base modification in analyzing nucleic acid molecules and acquiring data for analysis of nucleic acid molecules are described herein. Base modifications may include methylations. Methods to determine base modifications may include using features derived from sequencing. These features may include the pulse width of an optical signal from sequencing bases, the interpulse duration of bases, and the identity of the bases. Machine learning models can be trained to detect the base modifications using these features. The relative modification or methylation levels between haplotypes may indicate a disorder. Modification or methylation statuses may also be used to detect chimeric molecules.
Claims
exact text as granted — not AI-modified1 . A method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
(a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties:
for each nucleotide:
an identity of the nucleotide,
a position of the nucleotide within the sample nucleic acid molecule,
a width of the pulse corresponding to the nucleotide, and
an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide;
(b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the input data structure includes, for each nucleotide within the window, the properties:
the identity of the nucleotide,
a position of the nucleotide with respect to a target position within the window,
the width of the pulse corresponding to the nucleotide, and
the interpulse duration;
(c) inputting the input data structure into a model, the model trained by:
receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, each first data structure comprising values for the same properties as the input data structure,
storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and
optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and
(d) determining, using the model, whether the methylation is present inthe cytosine at the target position within the window in the input data structure.
2 . The method of claim 1 , wherein:
the input data structure is one input data structure of a plurality of input data structures, the sample nucleic acid molecule is one sample nucleic acid molecule of a plurality of sample nucleic acid molecules, the plurality of sample nucleic acid molecules is obtained from a biological sample of a subject, and each input data structure corresponds to a respective window of nucleotides sequenced in a respective sample nucleic acid molecule of the plurality of sample nucleic acid molecules, and the method further comprising:
receiving the plurality of input data structures,
inputting the plurality of input data structures into the model, and
determining, using the model, whether the methylation is present in a
cytosine at a target location in the respective window of each input data structure.
3 - 5 . (canceled)
6 . The method of claim 1 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression.
7 . The method of claim 1 , wherein:
the window of nucleotides corresponding to the input data structure comprises nucleotides on a first strand of the sample nucleic acid molecule and nucleotides on a second strand of the sample nucleic acid molecule, and the input data structure further comprises for each nucleotide within the window a value of a strand property, the strand property indicating the nucleotide being present on either the first strand or the second strand.
8 . The method of claim 1 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome.
9 - 10 . (canceled)
11 . The method of claim 1 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide.
12 . The method of claim 1 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule.
13 . (canceled)
14 . The method of claim 1 , further comprising:
validating the model using a plurality of nucleic acid molecules, each including a first portion corresponding to a first reference sequence and a second portion corresponding to a second reference sequence, wherein the first portion has a first methylation pattern, and the second portion has a second methylation pattern.
15 . The method of claim 14 , wherein the first portion is treated with a methylase.
16 . The method of claim 15 , wherein the second portion corresponds to an unmethylated portion of the second reference sequence.
17 - 20 . (canceled)
21 . A method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
(a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties:
for each nucleotide:
an identity of the nucleotide,
a position of the nucleotide within the sample nucleic acid molecule,
a width of the pulse corresponding to the nucleotide, and
an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide;
(b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the window comprises 6 consecutive nucleotides upstream of a nucleotide at a position within the window and 6 consecutive nucleotides downstream of the nucleotide at the target position, wherein the input data structure includes, for each nucleotide within the window, the properties:
the identity of the nucleotide,
a position of the nucleotide with respect to the target position ,
the width of the pulse corresponding to the nucleotide, and
the interpulse duration;
(c) inputting the input data structure into a model, the model trained by:
receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, the methylation being 5mC (5-methylcytosine), each first data structure comprising values for the same properties as the input data structure,
storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and
optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and
(d) determining, using the model, whether the 5mC methylation is present in the cytosine at the target position within the window in the input data structure.
22 . The method of claim 21 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression.
23 . A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that when executed control a computer system to perform a method for detecting a methylation of a cytosine in a nucleic acid molecule, the method comprising:
(a) receiving data acquired by sequencing a sample nucleic acid molecule by measuring pulses in an optical signal corresponding to nucleotides and obtaining, from the data, values for the following properties:
for each nucleotide:
an identity of the nucleotide,
a position of the nucleotide within the sample nucleic acid molecule,
a width of the pulse corresponding to the nucleotide, and
an interpulse duration representing a time between the pulse corresponding to the nucleotide and a pulse corresponding to a neighboring nucleotide;
(b) creating an input data structure, the input data structure comprising a window of the nucleotides sequenced in the sample nucleic acid molecule, wherein the input data structure includes, for each nucleotide within the window, the properties:
the identity of the nucleotide,
a position of the nucleotide with respect to a target position within the window,
the width of the pulse corresponding to the nucleotide, and
the interpulse duration;
(c) inputting the input data structure into a model, the model trained by:
receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective window of nucleotides sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in the optical signal corresponding to the nucleotides, wherein the methylation has a known first state in a cytosine at a target position in each window of each first nucleic acid molecule, each first data structure comprising values for the same properties as the input data structure,
storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the cytosine at the target position, and
optimizing, using the plurality of first training samples, parameters of the model based on outputs of the model matching or not matching corresponding labels of the first labels when the first plurality of first data structures is input to the model, wherein an output of the model specifies whether the cytosine at the target position in the respective window has the methylation; and
(d) determining, using the model, whether the methylation is present in the cytosine at the target position within the window in the input data structure.
24 . The computer product of claim 23 , wherein the methylation is 5mC (5-methylcytosine).
25 . The computer product of claim 23 , wherein:
the input data structure is one input data structure of a plurality of input data structures, the sample nucleic acid molecule is one sample nucleic acid molecule of a plurality of sample nucleic acid molecules, the plurality of sample nucleic acid molecules is obtained from a biological sample of a subject, and each input data structure corresponds to a respective window of nucleotides sequenced in a respective sample nucleic acid molecule of the plurality of sample nucleic acid molecules, and the method further comprising:
receiving the plurality of input data structures,
inputting the plurality of input data structures into the model, and
determining, using the model, whether the methylation is present in a cytosine at a target location in the respective window of each input data structure.
26 . (canceled)
27 . The computer product of claim 25 , wherein:
the plurality of sample nucleic acid molecules aligns to a plurality of genomic regions, for each genomic region of the plurality of genomic regions:
a number of sample nucleic acid molecules is aligned to the genomic region,
the number of sample nucleic acid molecules is greater than a cutoff number.
28 . The computer product of claim 23 , wherein the model includes a machine learning model, a principal component analysis, a convolutional neural network, or a logistic regression.
29 - 30 . (canceled)
31 . The method of claim 1 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position.
32 . The method of claim 1 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position.
33 . The method of claim 1 , wherein the window of the input data structure comprises 21 consecutive nucleotides upstream of the nucleotide at the target position and 21 consecutive nucleotides downstream of the nucleotide at the target position.
34 . The method of claim 21 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide.
35 . The computer product of claim 23 , wherein the optical signal is a fluorescence signal from a dye-labeled nucleotide.
36 . The method of claim 21 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome.
37 . The computer product of claim 23 , wherein the nucleotides within the window are determined using a circular consensus sequence and without alignment of the sequenced nucleotides to a reference genome.
38 . The method of claim 21 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule.
39 . The computer product of claim 23 , wherein each window associated with the first plurality of first data structures comprises 13 consecutive nucleotides on a first strand of each first nucleic acid molecule.
40 . The method of claim 21 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position.
41 . The computer product of claim 23 , wherein the window of the input data structure has a different number of consecutive nucleotides upstream of the nucleotide at the target position than the number of consecutive nucleotides downstream of the nucleotide at the target position.
42 . The method of claim 21 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position.
43 . The computer product of claim 23 , wherein the window of the input data structure comprises 10 consecutive nucleotides upstream of the nucleotide at the target position and 10 consecutive nucleotides downstream of the nucleotide at the target position.Join the waitlist — get patent alerts
Track US2023193360A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.