Methods and systems for genomic analysis
Abstract
A computer-implemented method for processing and/or analyzing nucleic acid sequencing data comprises receiving a first data input and a second data input. The first data input comprises untargeted sequencing data generated from a first nucleic acid sample obtained from a subject. The second data input comprises target-specific sequencing data generated from a second nucleic acid sample obtained from the subject. Next, with the aid of a computer processor, the first data input and the second data input are combined to produce a combined data set. Next, an output derived from the combined data set is generated. The output is indicative of the presence or absence of one or more polymorphisms of the first nucleic acid sample and/or the second nucleic acid sample.
Claims
exact text as granted — not AI-modified1 .- 28 . (canceled)
29 . A computer-implemented method, comprising:
(a) receiving, by a computing system, a set of data resulting from whole genome sequencing performed on a free DNA nucleic acid sample of a subject, wherein the free DNA nucleic acid sample is isolated from one or more of plasma or serum; (b) providing, by the computing system, as input to one or more models, one or more features of the set of data; (c) generating, by the computing system using the one or more models, output detecting one or more genomic regions within the set of data, wherein the one or more detected genomic regions comprise polymorphisms; and (d) generating, by the computing system, using an evaluation of the detected genomic regions, output indicating a likelihood of cancer tumor.
30 . The computer-implemented method of claim 29 , wherein the set of data further results from nucleic acid amplification prior to the whole genome sequencing.
31 . The computer-implemented method of claim 29 , further comprising:
selecting, by the computing system, the one or more features.
32 . The computer-implemented method of claim 31 , wherein the selecting the one or more features comprises the computing system using one or more of filter techniques, wrapper methods, embedded techniques, Benjamini-Hochberg procedures, Analysis of Variance (ANOVA), Wilcoxon approaches, or Threshold Number of Misclassification (TNoM).
33 . The computer-implemented method of claim 29 , wherein the one or more models comprise one or more of Hidden Markov Models (HMMs), random forest models, or Support Vector Machine (SVM) models.
34 . The computer-implemented method of claim 29 , wherein the generation of the output indicating the likelihood of cancer tumor comprises the computing system using one or more of HMMs, ANOVA, or Principal Component Analysis (PCA).
35 . The computer-implemented method of claim 29 , wherein said detected genomic regions comprise one or more of base changes, insertions, deletions, repeats, transversions, or copy number variants (CNVs).
36 . A system, comprising:
at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform: (a) receiving a set of data resulting from whole genome sequencing performed on a free DNA nucleic acid sample of a subject, wherein the free DNA nucleic acid sample is isolated from one or more of plasma or serum; (b) providing, as input to one or more models, one or more features of the set of data; (c) generating, using the one or more models, output detecting one or more genomic regions within the set of data, wherein the one or more detected genomic regions comprise polymorphisms; and (d) generating, using an evaluation of the detected genomic regions, output indicating a likelihood of cancer tumor.Join the waitlist — get patent alerts
Track US2022392577A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.