Artificial intelligence analysis of rna transcriptome for drug discovery
Abstract
A system and method may be provided to receive sample RNA reads from patients and generate lists of genes and their associated RNA expression levels in each patient. Some of the RNA reads may be matched to an RNA transcript or gene or gene family in terms of their match likelihood and other RNA reads may be matched to an RNA transcript or gene or gene family through the use of one or more machine learning classifiers. A machine learning classifier may be trained based on the plurality of the lists and a plurality of corresponding patients’ clinical status data to identify gene patterns that recur with a high degree of frequency in the plurality of the lists. Those gene patterns can be capable of modifying a disease or treatment response and can be targeted for drug/treatment development.
Claims
exact text as granted — not AI-modified1 - 32 . (canceled)
33 . A computer-implemented method comprising:
providing a list of sequenced RNA reads from a sample of a patient; providing a dictionary of RNA transcripts; indexing the list of RNA reads to the dictionary of RNA transcripts to determine for a first plurality of RNA reads a corresponding RNA transcript; determining for each of a second plurality of RNA reads that there is no corresponding RNA transcript in the dictionary of RNA transcripts; training a plurality of machine learning classifiers, to output a prediction for a member of a gene family, wherein each of the plurality of machine learning classifiers comprise at least a convolution neural network and/or a recurrent neural network; inputting each of the second plurality of the RNA reads into the plurality of machine learning classifiers, each of the plurality of machine learning classifiers outputting a prediction of whether the RNA read is a member of a gene family, and assigning two or more of the second plurality of the RNA reads to a highest probability gene family based on the outputs of the plurality of machine learning classifiers; assembling one or more RNA transcripts of the second plurality of RNA reads based on two or more of the second plurality of RNA reads mapping to the same gene family and having an overlapping sequence of nucleotides; generating a set of RNA transcript expression levels based on the corresponding RNA transcripts of the first plurality of RNA reads and the assembled RNA transcripts of the second plurality of RNA reads; creating a training set including the set of RNA transcript expression levels, an indication of whether the patient has a variation in treatment response (VITR), and data from other patients; training a response prediction machine learning classifier, based on the training set, to predict the existence of the VITR based on an input set of RNA transcript expression levels; and using one or more parameters of the response prediction machine learning classifier to identify at least one RNA transcript of interest or gene of interest in the VITR.Join the waitlist — get patent alerts
Track US2023238081A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.