US2021272695A1PendingUtilityA1

Systems and methods for using sequencing data for pathogen detection

Assignee: TEMPUS LABS INCPriority: Feb 26, 2019Filed: May 18, 2021Published: Sep 2, 2021
Est. expiryFeb 26, 2039(~12.6 yrs left)· nominal 20-yr term from priority
A61K 45/06G16H 10/40A61K 31/513C07K 16/22G16H 50/70G16H 20/10A61K 33/243G16H 50/20A61K 31/555Y02A50/30A61K 31/337Y02A90/10A61K 31/282
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods are provided for training a classifier to discriminate between a first cancer condition associated with an oncogenic pathogenic infection a second cancer condition that is not associated with an oncogenic pathogenic infection. Systems and methods are provided for distinguishing cancers associated with oncogenic pathogenic infections that contribute to the cancer pathology and cancers that are not associated with oncogenic pathogenic infections. Systems and methods are provided for treating cancer based on whether the cancer is associated with an oncogenic pathogenic infection.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for training a classifier to discriminate between a first cancer condition and a second cancer condition, wherein the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with an oncogenic pathogen free status, the method comprising:
 at a computer system comprising at least one processor and a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:   (A) obtaining a dataset comprising, for each respective subject in a plurality of subjects of a species: (i) a corresponding plurality of abundance values, wherein each respective abundance value in the corresponding plurality of abundance values quantifies a level of expression of a corresponding gene, in a plurality of genes, in a tumor sample of the respective subject, and (ii) an indication of cancer condition of the respective subject, wherein the indication of cancer condition identifies whether the respective subject has the first cancer condition or the second cancer condition, and wherein the plurality of subjects includes a first subset of subjects that are afflicted with the first cancer condition and a second subset of subjects that are afflicted with the second condition;   (B) identifying a discriminating gene set using the corresponding plurality of abundance values and respective indication of the cancer condition of respective subjects in the plurality of subjects, wherein the discriminating gene set comprises a subset of the plurality of genes; and   (C) using the respective abundance values for the discriminating gene set and the respective indication of cancer condition across the plurality of subjects to train a classifier to discriminate between the first cancer condition and the second cancer condition as a function of respective abundance values for the discriminating gene set.   
     
     
         2 . The method of  claim 1 , wherein the corresponding plurality of abundance values is obtained by RNA-seq. 
     
     
         3 . The method of  claim 1  or  2 , wherein each subject in the plurality of subjects afflicted with a first type of cancer and wherein the first type of cancer is one of breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophagus cancer, head/neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer. 
     
     
         4 . The method of  claim 1  or  2 , wherein each subject in the plurality of subjects afflicted with a first stage of a first type of cancer and wherein
 the first type of cancer is one of breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophagus cancer, head/neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer, and 
 the first stage of cancer is stage I, stage II, stage III, or stage IV. 
 
     
     
         5 . The method of any one of  claims 1 - 4 , wherein
 the plurality of subjects comprises one hundred subjects,   the first subset of subjects comprises twenty subjects, and   the second subset of subjects comprises twenty subjects.   
     
     
         6 . The method of any one of  claims 1 - 4 , wherein
 the plurality of subjects comprises one thousand subjects,   the first subset of subjects comprises one hundred subjects, and   the second subset of subjects comprises one hundred subjects.   
     
     
         7 . The method of any one of  claims 1 - 6 , wherein
 the species is human,   the plurality of genes comprises ten thousand genes, and   the discriminating gene set consists of between five and forty genes.   
     
     
         8 . The method of any one of  claims 1 - 6 , wherein
 the species is human,   the plurality of genes comprises five thousand genes, and   the discriminating gene set consists of between five and twenty-five genes.   
     
     
         9 . The method of  claim 1 , wherein the discriminating gene set consists of at least four-fold fewer genes than the plurality of genes. 
     
     
         10 . The method of any one of  claims 1 - 9 , wherein the identifying the discriminating gene set comprises:
 regressing the dataset based on all or a subset of the plurality of abundance values across the plurality of subjects against the respective indication of cancer condition across the plurality of subjects using a regression algorithm to thereby assign a corresponding regression coefficient, in a plurality of regression coefficients, to each respective gene in the plurality of genes, and   selecting those genes in the plurality of genes for the discriminating gene set that are assigned a coefficient by the regression algorithm that satisfies a coefficient threshold.   
     
     
         11 . The method of any one of  claims 1 - 9 , wherein the identifying the discriminating gene set comprises:
 splitting the dataset into a plurality of sets, wherein each set in the plurality of sets includes two or more subjects that are afflicted with the first cancer condition and two or more subjects that are afflicted with the second condition;   independently regressing each respective set in the plurality of sets based on all or a subset of the plurality of abundance values across the subjects of the respective set against the respective indication of cancer condition across the subject of the respective set using a regression algorithm to thereby assign a corresponding regression coefficient, in a plurality of regression coefficients, to each respective gene in the plurality of genes, and   selecting those genes in the plurality of genes for the discriminating gene set that are assigned a coefficient by the regression algorithm that satisfies a coefficient threshold for at least a threshold percentage of the plurality of sets.   
     
     
         12 . The method of  claim 11 , wherein the plurality of sets consists of between five and fifty sets. 
     
     
         13 . The method of  claim 11 , wherein the plurality of sets consists of ten sets. 
     
     
         14 . The method of  claim 10  or  11 , wherein the coefficient threshold is zero. 
     
     
         15 . The method of  claim 10  or  11 , wherein the coefficient threshold is satisfied when the absolute value of the corresponding regression coefficient is greater than zero. 
     
     
         16 . The method of  claim 10  or  11 , wherein the regression algorithm is logistic regression. 
     
     
         17 . The method of  claim 16 , wherein the logistic regression assumes: 
       
         
           
             
               
                 
                   P 
                   ⁡ 
                   
                     ( 
                     
                       Y 
                       = 
                       
                         1 
                         | 
                         
                           x 
                           i 
                         
                       
                     
                     ) 
                   
                 
                 = 
                 
                   
                     exp 
                     ⁡ 
                     
                       ( 
                       
                         
                           β 
                           0 
                         
                         + 
                         
                           
                             β 
                             1 
                           
                           ⁢ 
                           
                             x 
                             
                               i 
                               ⁢ 
                               1 
                             
                           
                         
                         + 
                         … 
                         + 
                         
                           
                             β 
                             k 
                           
                           ⁢ 
                           
                             x 
                             
                               i 
                               ⁢ 
                               k 
                             
                           
                         
                       
                       ) 
                     
                   
                   
                     1 
                     + 
                     
                       exp 
                       ⁡ 
                       
                         ( 
                         
                           
                             β 
                             0 
                           
                           + 
                           
                             
                               β 
                               1 
                             
                             ⁢ 
                             
                               x 
                               
                                 i 
                                 ⁢ 
                                 1 
                               
                             
                           
                           + 
                           … 
                           + 
                           
                             
                               β 
                               k 
                             
                             ⁢ 
                             
                               x 
                               
                                 i 
                                 ⁢ 
                                 k 
                               
                             
                           
                         
                         ) 
                       
                     
                   
                 
               
               , 
             
           
         
         wherein,
 x i =(x i1 , x i2 , . . . , x ik ) are the corresponding plurality of abundance values for the plurality of genes from the tumor sample of the i th  corresponding subject, 
 Y∈{0, 1} is a class label that has the value “1” when the corresponding subject i has the first cancer condition and has the value “0” when the corresponding subject i has the second cancer condition, wherein P(Y=1|x i ) is the estimated probability that the i th  corresponding subject is a member of the first cancer class, 
 β 0  is an intercept, 
 β j =(j=1, . . . k) is the plurality of regression coefficients, wherein each respective regression coefficient in the plurality of regression coefficients is for a corresponding gene in the plurality of genes, and wherein 
 the i th  corresponding subject is assigned to the first cancer class when P(Y=1|x i ) exceeds a predefined threshold value and to the second cancer class otherwise. 
 
       
     
     
         18 . The method of  claim 17 , wherein the predefined threshold value is 0.5. 
     
     
         19 . The method of  claim 17  or  18 , wherein the logistic regression is logistic least absolute shrinkage and selection operator (LASSO) regression in which β j  is subject to the constraint:
   min(Σ i=1   n [− y   i (β 0 +β 1   x   i + . . . +β k   x   ik )+log(1+exp(β 0 +β 1   x   i + . . . +β k   x   ik ))]),
 
 wherein,
 Σ j=1   k |β j |≤λ, and 
 λ is a constant. 
 
 
     
     
         20 . The method of  claim 10  or  11 , wherein the regression algorithm is logistic regression with L1 or L2 regularization. 
     
     
         21 . The method of any one of  claims 1 - 20 , wherein the species is human. 
     
     
         22 . The method of any one of  claims 1 - 21 , wherein the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. 
     
     
         23 . The method of any one of  claims 1 - 22 , wherein the at least one program further comprises instructions for:
 (D) using the classifier, subsequent to the using (C), to classify a test subject to the first cancer or to the second condition by inputting a test plurality of abundance values into the classifier, wherein each respective abundance value in the test plurality of abundance values quantifies a level of expression of a corresponding gene, in the plurality of genes, in a tumor sample of the test subject.   
     
     
         24 . The method of  claim 23 , further comprising:
 (E) providing a therapeutic intervention or imaging of the test subject based on a determination that the test subject has the first cancer condition or the second cancer condition.   
     
     
         25 . The method of any one of  claims 1 - 22 , wherein the at least one program further comprises instructions for:
 (D) using the classifier, subsequent to the using (C), to determine a likelihood that a test subject has the first cancer condition or a likelihood that the test subject has the second cancer condition by inputting a test plurality of abundance values into the classifier, wherein each respective abundance value in the test plurality abundance values quantifies a level of expression of a corresponding gene, in the plurality of genes, in a tumor sample of the test subject.   
     
     
         26 . The method of  claim 25 , further comprising:
 (E) providing a therapeutic intervention or imaging of the test subject based on the likelihood that the test subject has the first cancer condition or the second cancer condition.   
     
     
         27 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic virus. 
     
     
         28 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic virus listed in Table 1. 
     
     
         29 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic bacterium. 
     
     
         30 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic bacterium listed in Table 1. 
     
     
         31 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic trematode. 
     
     
         30 . The method of  claim 1 , wherein the first oncogenic pathogen is an oncogenic trematode listed in Table 1. 
     
     
         31 . The method of  claim 25 , wherein the classifier further uses one or more additional features of the test subject in addition to the test plurality of abundance values to classify the subject. 
     
     
         32 . The method of  claim 31 , wherein the one or more additional features comprises an amount of mutation of a predetermined gene in the test sample of the test subject. 
     
     
         33 . The method of  claim 31 , wherein the one or more additional features comprises an amount of mutation of each predetermined gene in a plurality of predetermined genes in the test sample of the test subject. 
     
     
         34 . A method for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with an HPV-free status, the method comprising:
 obtaining abundance data for at least five genes listed in Table 3 from a tumor sample of the human subject, and   inputting the abundance data into a classifier that is trained to discriminate between the first cancer condition and the second cancer condition, at least in part, based on the abundance of the at least five genes listed in Table 3.   
     
     
         35 . The method of  claim 34 , wherein the classifier is trained in accordance with any of the methods of  claims 1 - 25 . 
     
     
         36 . A plurality of nucleic acid probes for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with an HPV-free status, wherein:
 the plurality of nucleic acid probes comprises at least five nucleic acid probes, and   each of the at least five nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 3.   
     
     
         37 . The plurality of nucleic acid probes of  claim 36 , comprising at least ten nucleic acid probes, wherein each of the at least ten nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 3. 
     
     
         38 . The plurality of nucleic acid probes of  claim 36 , comprising at least twenty nucleic acid probes, wherein each of the at least twenty nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 3. 
     
     
         39 . The plurality of nucleic acid probes of  claim 36 , comprising at least twenty-four nucleic acid probes, wherein each of the at least twenty nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 3. 
     
     
         40 . The plurality of nucleic acid probes according to any one of  claims 36 - 39 , further comprising at least one nucleic acid probe comprising a nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a gene that is not listed in Table 3. 
     
     
         41 . The plurality of nucleic acid probes according to any one of  claims 36 - 40 , wherein each probe in the plurality of probes comprises a 5′ biotin-modified oligonucleotide. 
     
     
         42 . A method for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by an Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with an HPV-free status, the method comprising:
 obtaining abundance data for at least five genes listed in Table 4 from a tumor sample of the human subject, and   inputting the abundance data into a classifier that is trained to discriminate between the first cancer condition and the second cancer condition, at least in part, based on the abundance of the at least five genes listed in Table 4.   
     
     
         43 . The method of  claim 42 , wherein the classifier is trained in accordance with any of the methods of  claims 1 - 25 . 
     
     
         44 . A plurality of nucleic acid probes for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with an EBV-free status, wherein:
 the plurality of nucleic acid probes comprises at least five nucleic acid probes, and   each of the at least five nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 4.   
     
     
         45 . The plurality of nucleic acid probes of  claim 44 , comprising at least nine nucleic acid probes, wherein each of the at least ten nucleic acid probes comprises a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 4. 
     
     
         46 . The plurality of nucleic acid probes according to  claim 44  or  45 , further comprising at least one nucleic acid probe comprising a nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a gene that is not listed in Table 3. 
     
     
         47 . The plurality of nucleic acid probes according to any one of  claims 44 - 46 , wherein each probe in the plurality of probes comprises a 5′ biotin-modified oligonucleotide. 
     
     
         48 . A method for discriminating between a first cancer condition and a second cancer condition in a subject with a first type of cancer, wherein the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with an oncogenic pathogen-free status, the method comprising:
 (A) obtaining a dataset for the subject, the dataset comprising a plurality of abundance values, wherein each respective abundance value in the plurality of abundance values quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject; and   (B) inputting the dataset to a classifier trained according to the method of any one of  claims 1 - 25 .   
     
     
         49 . The method of  claim 48 , wherein the subject is afflicted with breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophagus cancer, head/neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer. 
     
     
         50 . The method of  claim 48 , wherein the dataset further comprises a variant allele count for one or more variant alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         51 . The method of  claim 50 , wherein the one or more variant alleles are selected from variant alleles in a gene selected from the group consisting of TP53 (ENSG00000141510), CDKN2A (ENSG00000147889), and PIK3CA (ENSG00000121879). 
     
     
         52 . The method according to any one of  claims 48 - 50 , wherein the first cancer condition is associated with infection by a first oncogenic pathogen selected from the group consisting of Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papilloma virus (HPV), human T-cell lymphotropic virus (HTLV-1), Kaposi's associated sarcoma virus (KSHV), and Merkel cell polyomavirus (MCV). 
     
     
         53 . The method according to any one of  claims 48 - 50 , wherein the first cancer condition is selected from the group consisting of cervical cancer associated with human papilloma virus (HPV), head and neck cancer associated with HPV, gastric cancer associated with Epstein-Barr virus (EBV), nasopharyngeal cancer associated with EBV, Burkitt lymphoma associated with EBV, Hodgkin lymphoma associated with EBV, liver cancer associated with hepatitis B virus (HBV), liver cancer associated with hepatitis C virus (HCV), Kaposi sarcoma associated with Kaposi's associated sarcoma virus (KSHV), adult T-cell leukemia/lymphoma associated with human T-cell lymphotropic virus (HTLV-1), and Merkel cell carcinoma associated with Merkel cell polyomavirus (MCV). 
     
     
         54 . A method for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with an HPV-free status, the method comprising:
 (A) obtaining a dataset for the subject, the dataset comprising a plurality of abundance values from the subject, wherein:
 each respective abundance value in the plurality abundance value quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject, and 
 the plurality of genes comprises at least five genes selected from the genes listed in Table 3; and 
   (B) inputting the dataset to a classifier trained to discriminate between at least the first cancer condition and the second cancer condition based on the abundance values of the plurality of genes.   
     
     
         55 . The method of  claim 54 , wherein the first cancer condition is cervical cancer associated with infection by a human papillomavirus (HPV). 
     
     
         56 . The method of  claim 54 , wherein the first cancer condition is head and neck cancer associated with infection by a human papillomavirus (HPV). 
     
     
         57 . The method according to any one of  claims 54 - 56 , wherein the plurality of genes comprises at least ten genes selected from the genes listed in Table 3. 
     
     
         58 . The method according to any one of  claims 54 - 56 , wherein the plurality of genes comprises at least twenty genes selected from the genes listed in Table 3. 
     
     
         59 . The method according to any one of  claims 54 - 56 , wherein the plurality of genes comprises at least all twenty-four of the genes listed in Table 3. 
     
     
         60 . The method according to any one of  claims 54 - 59 , wherein the plurality of genes comprises at least one gene that is not listed in Table 3. 
     
     
         61 . The method according to any one of  claims 54 - 60 , wherein the dataset further comprises a variant allele count for one or more alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         62 . The method of  claim 61 , wherein the one or more variant alleles are selected from variant alleles in a TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene. 
     
     
         63 . The method according to any one of  claims 54 - 62 , wherein the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. 
     
     
         64 . The method according to any one of  claims 54 - 62 , wherein the classifier was trained according to the method of any one of  claims 1 - 25 . 
     
     
         65 . A method for discriminating between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by an Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with an EBV-free status, the method comprising:
 (A) obtaining a dataset for the subject, the dataset comprising a plurality of abundance values from the subject, wherein:
 each respective abundance value in the plurality abundance value quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject, and 
 the plurality of genes comprises at least five genes selected from the genes listed in Table 4; and 
   (B) inputting the dataset to a classifier trained to discriminate between at least the first cancer condition and the second cancer condition based on the abundance values of the plurality of genes.   
     
     
         66 . The method of  claim 65 , wherein the first cancer condition is gastric cancer associated with infection by an Epstein-Barr virus (EBV). 
     
     
         67 . The method according to any one of  claims 65 - 66 , wherein the plurality of genes comprises at all nine genes listed in Table 4. 
     
     
         68 . The method according to any one of  claims 65 - 67 , wherein the plurality of genes comprises at least one gene that is not listed in Table 4. 
     
     
         69 . The method according to any one of  claims 65 - 68 , wherein the dataset further comprises a variant allele count for one or more alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         70 . The method of  claim 69 , wherein the one or more variant alleles are selected from variant alleles in a TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene. 
     
     
         71 . The method according to any one of  claims 65 - 70 , wherein the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. 
     
     
         72 . The method according to any one of  claims 65 - 71 , wherein the classifier was trained according to the method of any one of  claims 1 - 25 . 
     
     
         73 . A method for treating cervical cancer in a human cancer patient, the method comprising:
 (A) determining whether the human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by:
 obtaining a dataset for the human cancer patient, the dataset comprising a plurality of abundance values, wherein:
 each respective abundance value in the plurality abundance value quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject, and 
 the plurality of genes comprises at least five genes selected from the genes listed in Table 3, and 
 
 inputting the dataset to a classifier trained to discriminate between at least a first cancer condition associated with HPV infection and a second cancer condition associated with an HPV-free status based on the abundance values of the plurality of genes, in a cancerous tissue of the subject; and 
   (B) treating the cervical cancer by:
 when the classifier result indicates that the human cancer patient is infected with an HPV oncogenic virus, administering a first therapy tailored for treatment of cervical cancer associated with an HPV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an HPV oncogenic virus, administering a second therapy tailored for treatment of cervical cancer not associated with an HPV infection. 
   
     
     
         74 . The method of  claim 73 , wherein the plurality of genes comprises at least ten genes selected from the genes listed in Table 3. 
     
     
         75 . The method of  claim 73 , wherein the plurality of genes comprises at least twenty genes selected from the genes listed in Table 3. 
     
     
         76 . The method of  claim 73 , wherein the plurality of genes comprises at least all twenty-four of the genes listed in Table 3. 
     
     
         77 . The method according to any one of  claims 73 - 76 , wherein the plurality of genes comprises at least one gene that is not listed in Table 3. 
     
     
         78 . The method according to any one of  claims 73 - 77 , wherein the dataset further comprises a variant allele count for one or more alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         79 . The method of  claim 78 , wherein the one or more variant alleles are selected from variant alleles in a TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene. 
     
     
         80 . The method according to any one of  claims 73 - 79 , wherein the classifier a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. 
     
     
         81 . The method according to any one of  claims 73 - 80 , wherein the classifier was trained according to the method of any one of  claims 1 - 25 . 
     
     
         82 . The method according to any one of  claims 73 - 81 , wherein the first therapy tailored for treatment of cervical cancer associated with an HPV infection is a therapeutic vaccine. 
     
     
         83 . The method according to any one of  claims 73 - 81 , wherein the first therapy tailored for treatment of cervical cancer associated with an HPV infection is an adoptive cell therapy. 
     
     
         84 . The method according to any one of  claims 73 - 83 , wherein the second therapy tailored for treatment of cervical cancer not associated with an HPV infection is chemotherapy. 
     
     
         85 . The method of  claim 84 , wherein the chemotherapy includes administration of cisplatin. 
     
     
         86 . The method of  claim 85 , wherein the second therapy further comprises co-administration of a second therapeutic agent selected from the group consisting of 5-fluorouracil, paclitaxel, and bevacizumab. 
     
     
         87 . A method for treating head and neck cancer in a human cancer patient, the method comprising:
 (A) determining whether the human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by:
 obtaining a dataset for the human cancer patient, the dataset comprising a plurality of abundance values, wherein:
 each respective abundance value in the plurality abundance value quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject, and 
 the plurality of genes comprises at least five genes selected from the genes listed in Table 3, and 
 
 inputting the dataset to a classifier trained to discriminate between at least a first cancer condition associated with HPV infection and a second cancer condition associated with an HPV-free status based on the abundance values of the plurality of genes, in a cancerous tissue of the subject; and 
   (B) treating the head and neck cancer by:
 when the classifier result indicates that the human cancer patient is infected with an HPV oncogenic virus, administering a first therapy tailored for treatment of head and neck cancer associated with an HPV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an HPV oncogenic virus, administering a second therapy tailored for treatment of head and neck cancer not associated with an HPV infection. 
   
     
     
         88 . The method of  claim 87 , wherein the plurality of genes comprises at least ten genes selected from the genes listed in Table 3. 
     
     
         89 . The method of  claim 87 , wherein the plurality of genes comprises at least twenty genes selected from the genes listed in Table 3. 
     
     
         90 . The method of  claim 87 , wherein the plurality of genes comprises at least all twenty-four of the genes listed in Table 3. 
     
     
         91 . The method according to any one of  claims 87 - 90 , wherein the plurality of genes comprises at least one gene that is not listed in Table 3. 
     
     
         92 . The method according to any one of  claims 87 - 91 , wherein the dataset further comprises a variant allele count for one or more alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         93 . The method of  claim 92 , wherein the one or more variant alleles are selected from variant alleles in a TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene. 
     
     
         94 . The method according to any one of  claims 87 - 93 , wherein the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. 
     
     
         95 . The method according to any one of  claims 87 - 93 , wherein the classifier was trained according to the method of any one of  claims 1 - 25 . 
     
     
         96 . The method according to any one of  claims 87 - 95 , wherein the first therapy tailored for treatment of head and neck cancer associated with an HPV infection is a therapeutic vaccine. 
     
     
         97 . The method according to any one of  claims 87 - 95 , wherein the first therapy tailored for treatment of head and neck cancer associated with an HPV infection is an immune checkpoint inhibitor. 
     
     
         98 . The method according to any one of  claims 87 - 95 , wherein the first therapy tailored for treatment of head and neck cancer associated with an HPV infection is a PI3K inhibitor. 
     
     
         99 . The method according to any one of  claims 87 - 98 , wherein the second therapy tailored for treatment of head and neck cancer not associated with an HPV infection is chemotherapy. 
     
     
         100 . The method of  claim 99 , wherein the chemotherapy includes administration of cisplatin. 
     
     
         101 . The method of  claim 100 , wherein the second therapy further comprises concurrent radiotherapy or postoperative chemoradiation. 
     
     
         102 . A method for treating gastric cancer in a human cancer patient, the method comprising:
 (A) determining whether the human cancer patient is infected with an Epstein-Barr virus (EBV) oncogenic virus by:
 obtaining a dataset for the human cancer patient, the dataset comprising a plurality of abundance values, wherein:
 each respective abundance value in the plurality abundance value quantifies a level of expression of a corresponding gene, in a plurality of genes, in a cancerous tissue from the subject, and 
 the plurality of genes comprises at least five genes selected from the genes listed in Table 4, and 
 
 inputting the dataset to a classifier trained to discriminate between at least a first cancer condition associated with EBV infection and a second cancer condition associated with an EBV-free status based on the abundance values of the plurality of genes, in a cancerous tissue of the subject; and 
   (B) treating the gastric cancer by:
 when the classifier result indicates that the human cancer patient is infected with an EBV oncogenic virus, administering a first therapy tailored for treatment of gastric cancer associated with an EBV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an EBV oncogenic virus, administering a second therapy tailored for treatment of gastric cancer not associated with an EBV infection. 
   
     
     
         103 . The method of  claim 102 , wherein the plurality of genes comprises at least all nine of the genes listed in Table 4. 
     
     
         104 . The method according to any one of  claims 102 - 103 , wherein the plurality of genes comprises at least one gene that is not listed in Table 4. 
     
     
         105 . The method according to any one of  claims 102 - 104 , wherein the dataset further comprises a variant allele count for one or more alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         106 . The method of  claim 105 , wherein the one or more variant alleles are selected from variant alleles in a TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene. 
     
     
         107 . The method according to any one of  claims 102 - 106 , wherein the classifier is a multivariate logistic regression algorithm, a neural network algorithm, or a convolutional neural network algorithm. 
     
     
         108 . The method according to any one of  claims 102 - 106 , wherein the classifier is a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm. 
     
     
         109 . The method according to any one of  claims 102 - 108 , wherein the classifier was trained according to the method of any one of  claims 1 - 25 . 
     
     
         110 . The method according to any one of  claims 102 - 109 , wherein the first therapy tailored for treatment of gastric cancer associated with an EBV infection is an immune checkpoint inhibitor. 
     
     
         111 . The method according to any one of  claims 102 - 110 , wherein the second therapy tailored for treatment of gastric cancer not associated with an EBV infection is chemotherapy. 
     
     
         112 . The method according to  claim 111 , wherein the chemotherapy includes administration of a therapeutic agent selected from the group consisting of paclitaxel, carboplatin, cisplatin, 5-fluorouracil, and oxaliplatin. 
     
     
         113 . The method according to  claim 111 , wherein the chemotherapy includes administration of paclitaxel and carboplatin. 
     
     
         114 . The method according to  claim 111 , wherein the chemotherapy includes administration of cisplatin and 5-fluorouracil. 
     
     
         115 . The method according to  claim 111 , wherein the chemotherapy includes administration of oxaliplatin and 5-fluorouracil. 
     
     
         116 . The method according to any one of  claims 48  to  115 , further comprising:
 determining the plurality of abundance values by RNA sequencing of a sample of the cancerous tissue from the human cancer patient. 
 The method according to any one of  claims 48  to  115 , wherein the cancerous tissue from the subject is a tumor sample from the subject. 
 
     
     
         117 . An electronic device, comprising:
 one or more processors;   memory; and   one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods of  claims 48  to  116 .   
     
     
         118 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by an electronic device with one or more processors and a memory cause the device to perform any of the methods of  claims 48  to  116 . 
     
     
         119 . A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for training a classifier to discriminate between a first cancer condition and a second cancer condition, wherein the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with an oncogenic pathogen-free status, the method comprising:
 (A) obtaining a dataset comprising, for each respective subject in a plurality of subjects of a species: (i) a corresponding plurality of abundance values, wherein each respective abundance value in the corresponding plurality of abundance values quantifies a level of expression of a corresponding gene, in a plurality of genes, in a tumor sample of the respective subject, and (ii) an indication of cancer condition of the respective subject, wherein the indication of cancer condition identifies whether the respective subject has the first cancer condition or the second cancer condition, and wherein the plurality of subjects includes a first subset of subjects that are afflicted with the first cancer condition and a second subset of subjects that are afflicted with the second condition;   (B) identifying a discriminating gene set using the corresponding plurality of abundance values and respective indication of the cancer condition of respective subjects in the plurality of subjects, wherein the discriminating gene set comprises a subset of the plurality of genes; and   (C) using the respective abundance values for the discriminating gene set and the respective indication of cancer condition across the plurality of subjects to train a classifier to discriminate between the first cancer condition and the second cancer condition as a function of respective abundance values for the discriminating gene set.   
     
     
         120 . A computer system for training a classifier to discriminate between a first cancer condition and a second cancer condition, wherein the first cancer condition is associated with infection by a first oncogenic virus and the second cancer condition is associated with an oncogenic viral free status, the computer system comprising:
 at least one processor, and   a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:   (A) obtaining a dataset comprising, for each respective subject in a plurality of subjects of a species: (i) a corresponding plurality of abundance values, wherein each respective abundance value in the corresponding plurality of abundance values quantifies a level of expression of a corresponding gene, in a plurality of genes, in a tumor sample of the respective subject, and (ii) an indication of cancer condition of the respective subject, wherein the indication of cancer condition identifies whether the respective subject has the first cancer condition or the second cancer condition, and wherein the plurality of subjects includes a first subset of subjects that are afflicted with the first cancer condition and a second subset of subjects that are afflicted with the second condition;   (B) identifying a discriminating gene set using the corresponding plurality of abundance values and respective indication of the cancer condition of respective subjects in the plurality of subjects, wherein the discriminating gene set comprises a subset of the plurality of genes; and   (C) using the respective abundance values for the discriminating gene set and the respective indication of cancer condition across the plurality of subjects to train a classifier to discriminate between the first cancer condition and the second cancer condition as a function of respective abundance values for the discriminating gene set.   
     
     
         122 . A method for discriminating between a first cancer condition and a second cancer condition in a subject with a first type of cancer, wherein the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with an oncogenic pathogen-free status, the method comprising:
 at a computer system comprising at least one processor and a memory storing at least one program for execution by the at least one processor
 (A) obtaining a dataset for the subject, the dataset comprising a plurality of abundance values, wherein each respective abundance value in the plurality of abundance values quantifies a level of expression of a corresponding gene, in a discriminating gene set, in a cancerous tissue from the subject; and 
 (B) inputting the dataset to a classifier trained to discriminate between at least the first cancer condition and the second cancer condition based on abundance values for the discriminating gene set in a cancerous tissue of a subject, thereby determining the cancer condition of the subject. 
   
     
     
         123 . The method of  claim 122 , wherein the first type of cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophagus cancer, head/neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer. 
     
     
         124 . The method of  claim 122 , wherein the dataset further comprises a variant allele count for one or more variant alleles at one or more locus in the genome of the cancerous tissue from the subject. 
     
     
         125 . The method of  claim 122 , wherein the first cancer condition is associated with infection by a first oncogenic pathogen selected from the group consisting of Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papilloma virus (HPV), human T-cell lymphotropic virus (HTLV-1), Kaposi's associated sarcoma virus (KSHV), and Merkel cell polyomavirus (MCV). 
     
     
         126 . The method of  claim 122 , wherein the first cancer condition is selected from the group consisting of cervical cancer associated with human papilloma virus (HPV), head and neck cancer associated with HPV, gastric cancer associated with Epstein-Barr virus (EBV), nasopharyngeal cancer associated with EBV, Burkitt lymphoma associated with EBV, Hodgkin lymphoma associated with EBV, liver cancer associated with hepatitis B virus (HBV), liver cancer associated with hepatitis C virus (HCV), Kaposi sarcoma associated with Kaposi's associated sarcoma virus (KSHV), adult T-cell leukemia/lymphoma associated with human T-cell lymphotropic virus (HTLV-1), and Merkel cell carcinoma associated with Merkel cell polyomavirus (MCV). 
     
     
         127 . The method of  claim 122 , wherein:
 the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with an HPV-free status, and   the discriminating gene set comprises at least five genes selected from the genes listed in Table 3.   
     
     
         128 . The method of  claim 127 , wherein the first cancer condition is cervical cancer associated with infection by a human papillomavirus (HPV). 
     
     
         129 . The method of  claim 127 , wherein the first cancer condition is head and neck cancer associated with infection by a human papillomavirus (HPV). 
     
     
         130 . The method of  claim 127 , wherein the discriminating gene set comprises at least ten genes selected from the genes listed in Table 3. 
     
     
         131 . The method of  claim 127 , wherein the discriminating gene set comprises at least twenty genes selected from the genes listed in Table 3. 
     
     
         132 . The method of  claim 127 , wherein the discriminating gene set comprises at least all twenty-four of the genes listed in Table 3. 
     
     
         133 . The method of  claim 127 , wherein the dataset further comprises a variant allele count for TP53 (ENSG00000141510) and CDKN2A (ENSG00000147889) in the genome of the cancerous tissue from the subject. 
     
     
         134 . The method of  claim 128 , the method further comprising:
 (C) treating the subject for cervical cancer by:
 when the classifier result indicates that the human cancer patient is infected with an HPV oncogenic virus, administering a first therapy tailored for treatment of cervical cancer associated with an HPV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an HPV oncogenic virus, administering a second therapy tailored for treatment of cervical cancer not associated with an HPV infection. 
   
     
     
         135 . The method of  claim 134 , wherein the first therapy tailored for treatment of cervical cancer associated with an HPV infection comprises a therapeutic vaccine or an adoptive cell therapy. 
     
     
         136 . The method of  claim 134 , wherein the second therapy tailored for treatment of cervical cancer not associated with an HPV infection is chemotherapy. 
     
     
         137 . The method of  claim 136 , wherein the chemotherapy comprises co-administration of cisplatin and a second therapeutic agent selected from the group consisting of 5-fluorouracil, paclitaxel, and bevacizumab. 
     
     
         138 . The method of  claim 129 , the method further comprising:
 (C) treating the subject for head and neck cancer by:
 when the classifier result indicates that the human cancer patient is infected with an HPV oncogenic virus, administering a first therapy tailored for treatment of head and neck cancer associated with an HPV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an HPV oncogenic virus, administering a second therapy tailored for treatment of head and neck cancer not associated with an HPV infection. 
   
     
     
         139 . The method of  claim 138 , wherein the first therapy tailored for treatment of head and neck cancer associated with an HPV infection comprises a therapeutic vaccine, an immune checkpoint inhibitor, or a PI3K inhibitor. 
     
     
         140 . The method of  claim 138 , wherein the second therapy tailored for treatment of head and neck cancer not associated with an HPV infection is chemotherapy. 
     
     
         141 . The method of  claim 99 , wherein:
 the chemotherapy includes administration of cisplatin, and   the second therapy further comprises concurrent radiotherapy or postoperative chemoradiation.   
     
     
         142 . The method of  claim 122 , wherein:
 the first cancer condition is associated with infection by an Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with an EBV-free status, and   the discriminating gene set comprises at least five genes selected from the genes listed in Table 4.   
     
     
         143 . The method of  claim 134 , wherein the first cancer condition is gastric cancer associated with infection by an Epstein-Barr virus (EBV). 
     
     
         144 . The method of  claim 134 , wherein the discriminating gene set comprises all nine genes listed in Table 4. 
     
     
         145 . The method of  claim 134 , wherein the dataset further comprises a variant allele count for TP53 (ENSG00000141510) and PIK3CA (ENSG00000121879) in the genome of the cancerous tissue from the subject. 
     
     
         146 . The method of  claim 142 , the method further comprising:
 (C) treating the subject for gastric cancer by:
 when the classifier result indicates that the human cancer patient is infected with an EBV oncogenic virus, administering a first therapy tailored for treatment of gastric cancer associated with an EBV infection, and 
 when the classifier result indicates that the human cancer patient is not infected with an EBV oncogenic virus, administering a second therapy tailored for treatment of gastric cancer not associated with an EBV infection. 
   
     
     
         147 . The method of  claim 146 , wherein the first therapy tailored for treatment of gastric cancer associated with an EBV infection comprises an immune checkpoint inhibitor. 
     
     
         148 . The method of  claim 146 , wherein the second therapy tailored for treatment of gastric cancer not associated with an EBV infection is chemotherapy. 
     
     
         149 . The method according to  claim 148 , wherein the chemotherapy includes administration of a therapeutic agent selected from the group consisting of paclitaxel, carboplatin, cisplatin, 5-fluorouracil, and oxaliplatin. 
     
     
         150 . The method of  claim 122 , the method further comprising:
 (C) treating the subject for cancer by:
 when the classifier result indicates that the human cancer patient is infected with the first oncogenic pathogen, administering a first therapy tailored for treatment of the first type of cancer associated with infection by the first oncogenic pathogen, and 
 when the classifier result indicates that the human cancer patient is not infected with the first oncogenic pathogen, administering a second therapy tailored for treatment of the first type of cancer associated with an oncogenic pathogen-free status. 
   
     
     
         151 . The method of  claim 122 , wherein the classifier was trained by a method comprising:
 (1) obtaining a dataset comprising, for each respective subject in a plurality of subjects of a species: (i) a corresponding plurality of abundance values, wherein each respective abundance value in the corresponding plurality of abundance values quantifies a level of expression of a corresponding gene, in a plurality of genes, in a tumor sample of the respective subject, and (ii) an indication of cancer condition of the respective subject, wherein the indication of cancer condition identifies whether the respective subject has the first cancer condition or the second cancer condition, and wherein the plurality of subjects includes a first subset of subjects that are afflicted with the first cancer condition and a second subset of subjects that are afflicted with the second condition;   (2) identifying the discriminating gene set using the corresponding plurality of abundance values and respective indication of the cancer condition of respective subjects in the plurality of subjects, wherein the discriminating gene set comprises a subset of the plurality of genes; and   (3) using the respective abundance values for the discriminating gene set and the respective indication of cancer condition across the plurality of subjects to train a classifier to discriminate between the first cancer condition and the second cancer condition as a function of respective abundance values for the discriminating gene set.

Join the waitlist — get patent alerts

Track US2021272695A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.