US2023055403A1PendingUtilityA1

Methods and systems for multiple taxonomic classification

Assignee: UNIV UTAH RES FOUNDPriority: Apr 24, 2015Filed: Apr 19, 2022Published: Feb 23, 2023
Est. expiryApr 24, 2035(~8.7 yrs left)· nominal 20-yr term from priority
G16B 40/20G16B 30/10G16B 40/00Y02A90/10C40B 20/00G16B 30/00
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Described herein are methods of identifying a plurality of polynucleotides, as well as detecting presence, absence, or abundance of a plurality of taxa in a sample. Also provided are systems for performing methods of the disclosure.

Claims

exact text as granted — not AI-modified
1 .- 46 . (canceled) 
     
     
         47 . A method of detecting a plurality of taxa in a sample, the method comprising providing sequencing reads for a plurality of polynucleotides from the sample, and for each sequencing read:
 (a) assigning the sequencing read to a first taxonomic group based on a first sequence comparison between the sequencing read and a first plurality of polynucleotide sequences from the different first taxonomic groups, wherein at least two sequencing reads are assigned to different taxonomic groups;   (b) performing with a computer system a second sequence comparison between the sequencing read and a second plurality of polynucleotide sequences corresponding to members of the first taxonomic group, wherein the comparison comprises counting a number of k-mers within the sequencing read of at least 5 nucleotides in length that exactly match one or more k-mers within a reference sequence in the second plurality of polynucleotide sequences;   (c) classifying the sequencing read as belonging to a second taxonomic group that is more specific than the first taxonomic group if a measure of similarity between the sequencing read and reference sequence is above a first threshold level;   (d) if no similarity above the first threshold level is identified in (c), classifying the sequencing read as belonging to the second taxonomic group based on similarity above a second threshold level determined by comparing with the computer system a sequence derived from translating the sequencing read and a third set of reference sequences corresponding to amino acid sequences of members of the first taxonomic group; and   (e) identifying the presence, absence, or abundance of the plurality of taxa in the sample based on the classifying of the sequencing reads.   
     
     
         48 . The method of  claim 47 , wherein step (b) further comprises calculating k-mer weights as measures of how likely it is that k-mers within the sequencing read are derived from a reference sequence in the second plurality of polynucleotide sequences. 
     
     
         49 . The method of  claim 47 , wherein the third set of reference sequences consist of polynucleotide sequences derived from reverse-translating the corresponding amino acid sequences. 
     
     
         50 . The method of  claim 47 , further comprising performing with the computer system a relaxed sequence comparison between the sequencing read and the second plurality of polynucleotide sequences if the similarity in (d) is below the second threshold, wherein the relaxed sequence comparison is less stringent than the second sequence comparison. 
     
     
         51 . The method of  claim 47 , wherein classifying the sequencing read in step (c) comprises resolving a tie between two or more possible taxonomic groups based on a k-mer weight as a measure of how likely it is that the sequencing read corresponds to a polynucleotide from an ancestor of one of the possible taxonomic groups. 
     
     
         52 . The method of  claim 47 , further comprising diagnosing a condition based on a degree of similarity between the plurality of taxa detected in the sample and a biological signature for the condition. 
     
     
         53 . The method of  claim 52 , wherein the condition is contamination of the sample. 
     
     
         54 . The method of  claim 52 , wherein the condition is an infection of a subject. 
     
     
         55 . The method of  claim 54 , wherein infection is assessed based on the presence or amount of (i) sequences of host transcripts; and/or (ii) sequences of one or more infectious agents. 
     
     
         56 . The method of  claim 54 , further comprising monitoring treatment in an infected subject by detecting presence, absence, or abundance of a plurality of taxa in samples from the infected subject at multiple times after beginning treatment. 
     
     
         57 . The method of  claim 56 , further comprising changing treatment of the infected subject based on results of the monitoring. 
     
     
         58 . The method of  claim 47 , wherein step (c) further comprises classifying the sequencing read as corresponding to a gene transcript if the measure of similarity between the sequencing read and reference sequence is above the first threshold level. 
     
     
         59 . The method of  claim 58 , further comprising diagnosing a condition based on a degree of similarity between the plurality of taxa detected in the sample and a biological signature for the condition. 
     
     
         60 . The method of  claim 47 , wherein step (a) comprises assigning sequencing reads to two or more taxa selected from bacteria, viruses, fungi, or humans. 
     
     
         61 . The method of  claim 47 , wherein a sequencing read classified as belonging to the second taxonomic group and not present among the group of sequences corresponding to the second taxonomic group is added to the group of sequences corresponding to the second taxonomic group for use in later sequence comparisons. 
     
     
         62 . The method of  claim 47 , wherein the second plurality of nucleotide sequences comprises marker gene sequences for taxonomic classification of bacterial sequences. 
     
     
         63 . The method of  claim 62 , wherein the marker gene sequences comprise 16S rRNA sequences. 
     
     
         64 . The method of  claim 47 , wherein the second plurality of nucleotide sequences comprises sequences of human transcripts. 
     
     
         65 .- 69 . (canceled)

Join the waitlist — get patent alerts

Track US2023055403A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.