Machine Learning for Somatic Single Nucleotide Variant Detection in Cell-free Tumor Nucleic acid Sequencing Applications
Abstract
Systems and methods are disclosed to detect single-nucleotide variations (SNVs) from somatic sources in a cell-free biological sample of a subject by generating training data with class labels; in computer memory, generating a machine learning unit comprising one output for each of adenine (A), cytosine (C), guanine (G), and thymine (T) calls; training the machine learning unit; and applying the machine learning unit to detect the SNVs from somatic sources in the cell-free biological sample of the subject, wherein the cell-free biological sample comprises a mixture of nucleic acid molecules from somatic and germline sources.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
(a) providing a training data set comprising, for each mixture in a plurality of mixtures, wherein each mixture in the plurality comprises polynucleotides from a plurality of different subjects, values indicating: (i) a quantitative measure of each of a plurality of bases at each of a plurality of genomic base positions from sequence reads of a plurality of polynucleotides in the mixture, and (ii) a plurality of class labels, each class label classifying the mixture as having one or more particular bases at a particular genomic base position; and (b) training a machine learning unit on the training data set to generate one or more classification models for detecting a presence of a base at each of a plurality of genomic base positions in a test sample.
2 . The method of claim 1 , wherein the plurality of mixtures is provided by combining polynucleotides from a plurality of samples from different subjects in predetermined amounts.
3 . The method of claim 1 , wherein the one or more classification models further call a relative frequency of each of one or more bases at each of a plurality of genomic base positions in the test sample.
4 . The method of claim 1 , wherein the one or more classification models further call a probability of the presence of each of one or more bases at each of a plurality of genomic base positions in the test sample.
5 . The method of claim 1 , wherein the machine learning unit comprises a four-output neural network, a three-output neural network, a support vector machine (SVM), or another supervised learning machine.
6 . The method of claim 1 , further comprising:
(c) providing a test data set comprising, for a test sample, values indicating a quantitative measure of each of a plurality of bases at each of a plurality of genomic base positions for sequence reads of a plurality of polynucleotides in the test sample; and (d) using a classification model of (b) to call the presence of a base at a plurality of the genomic base positions.
7 . A training set, comprising a plurality of mixtures of cell-free DNA, wherein each mixture comprises a first normal cell-free DNA sample in a second normal cell-free DNA sample in various predetermined permutations and relative concentrations.Join the waitlist — get patent alerts
Track US2022199197A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.