US2019367995A1PendingUtilityA1

Biomarkers for colorectal cancer

Assignee: BGI SHENZHEN CO LTDPriority: Aug 6, 2013Filed: Aug 15, 2019Published: Dec 5, 2019
Est. expiryAug 6, 2033(~7 yrs left)· nominal 20-yr term from priority
C12Q 1/6886C12Q 2600/16C12Q 2600/112C12Q 2600/158C12Q 1/6806
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Biomarkers and methods for predicting the risk of a disease related to microbiota, in particular colorectal cancer (CRC), are described.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 1) obtaining sequencing reads from sample j of a subject, wherein the sample j comprises microbiota;   2) mapping the sequencing reads to a gene catalog and deriving a gene profile from the mapping result;   3) determining the relative abundance of each gene marker in a set of gene markers comprising at least three genes having the nucleotide sequences of SEQ ID NO: 10, SEQ ID NO: 14 and SEQ ID NO: 6; and   4) calculating an index of sample j using the following formula:   
       
         
           
             
               
                 
                   I 
                   j 
                 
                 = 
                 
                   [ 
                   
                     
                       
                         
                           ∑ 
                           i 
                         
                          
                         
                           
                             ϵ 
                             N 
                           
                            
                           log 
                            
                           
                               
                           
                            
                           10 
                            
                           
                             ( 
                             
                               
                                 A 
                                 ij 
                               
                               + 
                               
                                 10 
                                 
                                   - 
                                   20 
                                 
                               
                             
                             ) 
                           
                         
                       
                       
                          
                         N 
                          
                       
                     
                     - 
                     
                       
                         
                           ∑ 
                           i 
                         
                          
                         
                           
                             ϵ 
                             M 
                           
                            
                           log 
                            
                           
                               
                           
                            
                           10 
                            
                           
                             ( 
                             
                               
                                 A 
                                 ij 
                               
                               + 
                               
                                 10 
                                 
                                   - 
                                   20 
                                 
                               
                             
                             ) 
                           
                         
                       
                       
                          
                         M 
                          
                       
                     
                   
                   ] 
                 
               
               , 
             
           
         
       
       wherein:
 A ij  is the relative abundance of marker i in sample j, wherein i refers to each of the gene markers in the gene marker set, 
 N is a subset of all of the abnormal-associated gene markers in selected biomarkers related to the abnormal condition, 
 M is a subset of all of the control-associated gene markers in selected biomarkers related to the abnormal condition, and 
 |N| and |M| are numbers (sizes) of the biomarkers in these two subsets, respectively, 
 5) identifying the subject as having or being at a risk of developing the abnormal condition when the index is greater than a cutoff, and 
 6) modifying the risk by altering metabolic, immunological, or developmental pathways in subject. 
 
     
     
         2 . The method of  claim 1 , wherein the method further comprises estimating the false discovery rate (FDR). 
     
     
         3 . The method of  claim 1 , wherein the gene catalog is a non-redundant gene set constructed for related microbiota, and the set of gene markers further comprises one or more genes having the nucleotide sequences of SEQ ID NOs: 1 to 5, SEQ ID NOs: 7 to 9, SEQ ID NOs: 11 to 13, and SEQ ID NOs: 15 to 31. 
     
     
         4 . The method of  claim 1 , wherein the abnormal condition related to microbiota is an abnormal condition related to environmental microbiota. 
     
     
         5 . The method of  claim 1 , wherein the abnormal condition related to microbiota is a disease related to microbiota present in the animal body or the human body, wherein the microbiota is selected from the group consisting of microbiota found in the gastrointestinal tract, nasal passages, oral cavities, skin and the urogenital tract. 
     
     
         6 . The method of  claim 1 , wherein the abnormal condition related to microbiota is a colorectal disease selected from the group consisting of Colorectal Cancer, Ulcerative Colitis, Crohn's Disease, Irritable Bowel Syndrome (IBS), Diverticular Disease, Hemorrhoids, Anal Fissure, and Bowel Incontinence. 
     
     
         7 . The method of  claim 1 , wherein the sequencing reads are obtained via a next-generation sequencing method or a next-next-generation sequencing method. 
     
     
         8 . The method of  claim 1 , wherein the cutoff value is obtained by a Receiver Operator Characteristic (ROC) method, wherein the cutoff corresponds to the value when the AUC (Area Under the Curve) is at its maximum. 
     
     
         9 . The method of  claim 1 , wherein the sample is a feces sample, a nasal cavity swab, a buccal swab, a skin swab or a vaginal swab. 
     
     
         10 . The method of  claim 1 , wherein the sequencing reads are obtained via steps comprising:
 1) collecting the sample j from the subject;   2) extracting DNA from the sample;   3) constructing a DNA library; and   4) sequencing the library.   
     
     
         11 . A method, comprising:
 1) obtaining sequencing reads from sample j of the subject, wherein the sample j comprises microbiota;   2) mapping the sequencing reads to a human gut gene catalog and deriving a gene profile from the mapping result;   3) determining the relative abundance of each of the gene markers listed in SEQ ID NOs: 1-31; and   4) calculating an index of sample j using the following formula:   
       
         
           
             
               
                 
                   I 
                   j 
                 
                 = 
                 
                   [ 
                   
                     
                       
                         
                           ∑ 
                           i 
                         
                          
                         
                           
                             ϵ 
                             N 
                           
                            
                           log 
                            
                           
                               
                           
                            
                           10 
                            
                           
                             ( 
                             
                               
                                 A 
                                 ij 
                               
                               + 
                               
                                 10 
                                 
                                   - 
                                   20 
                                 
                               
                             
                             ) 
                           
                         
                       
                       
                          
                         N 
                          
                       
                     
                     - 
                     
                       
                         
                           ∑ 
                           i 
                         
                          
                         
                           
                             ϵ 
                             M 
                           
                            
                           log 
                            
                           
                               
                           
                            
                           10 
                            
                           
                             ( 
                             
                               
                                 A 
                                 ij 
                               
                               + 
                               
                                 10 
                                 
                                   - 
                                   20 
                                 
                               
                             
                             ) 
                           
                         
                       
                       
                          
                         M 
                          
                       
                     
                   
                   ] 
                 
               
               , 
             
           
         
       
       wherein: 
       A ij  is the relative abundance of marker i in sample j, wherein i refers to each of the gene markers listed in SEQ ID NOs 1-31, 
       N is a subset of all of colorectal cancer (CRC)-associated gene markers and M is a subset of all of the control-associated gene markers, 
       wherein the subset of CRC-associated gene markers and the subset of control-associated gene markers are shown in Table 1, and 
       |N| and |M| are numbers (sizes) of the biomarkers in these two subsets, respectively,
 5) identifying the subject as having or being at a risk of developing CRC when the index is greater than a cutoff, and 
 6) modifying the risk by altering metabolic, immunological, or developmental pathways in subject. 
 
     
     
         12 . The method of  claim 11 , wherein the cutoff value is obtained by a Receiver Operator Characteristic (ROC) method, wherein the cutoff corresponds to the value when the AUC (Area Under the Curve) is at its maximum. 
     
     
         13 . The method of  claim 12 , wherein the value of the cutoff is −0.0575. 
     
     
         14 . The method of  claim 11 , wherein the sequencing reads are obtained via steps comprising:
 1) collecting the sample j from the subject;   2) extracting DNA from the sample;   3) constructing a DNA library; and   4) sequencing the library.   
     
     
         15 . A method of diagnosing whether a subject has colorectal cancer or is at the risk of developing colorectal cancer (CRC), comprising:
 1) obtaining a feces sample j from the subject;   2) measuring the abundance information of each marker gene in a gene marker set comprising at least two genes selected from the group consisting of SEQ ID NOs: 1 to 31 in sample j using quantitative PCR;   3) calculating an index of sample j using the following formula:   
       
         
           
             
               
                 I 
                 j 
               
               = 
               
                 [ 
                 
                   
                     
                       
                         ∑ 
                         i 
                       
                        
                       
                         
                           ϵ 
                           N 
                         
                          
                         log 
                          
                         
                             
                         
                          
                         10 
                          
                         
                           ( 
                           
                             
                               A 
                               ij 
                             
                             + 
                             
                               10 
                               
                                 - 
                                 20 
                               
                             
                           
                           ) 
                         
                       
                     
                     
                        
                       N 
                        
                     
                   
                   - 
                   
                     
                       
                         ∑ 
                         i 
                       
                        
                       
                         
                           ϵ 
                           M 
                         
                          
                         log 
                          
                         
                             
                         
                          
                         10 
                          
                         
                           ( 
                           
                             
                               A 
                               ij 
                             
                             + 
                             
                               10 
                               
                                 - 
                                 20 
                               
                             
                           
                           ) 
                         
                       
                     
                     
                        
                       M 
                        
                     
                   
                 
                 ] 
               
             
           
         
       
       wherein:
 A ij  is the relative abundance of marker i in sample j, wherein i refers to each of the gene markers of the gene marker set, 
 N is a subset of all of colorectal cancer (CRC)-associated gene markers and M is a subset of all of the control-associated gene markers, 
 wherein the subset of CRC-associated gene markers and the subset of control-associated gene markers are shown in Table 1, and 
 |N| and |M| are numbers (sizes) of the biomarkers in these two subsets, respectively, 
 4) identifying the subject as having or being at a risk of developing CRC when the index is greater than a cutoff, and 
 5) modifying the risk by altering metabolic, immunological, or developmental pathways in subject. 
 
     
     
         16 . The method of  claim 15 , wherein the cutoff value is obtained by a Receiver Operator Characteristic (ROC) method, wherein the cutoff corresponds to the value when the AUC (Area Under the Curve) is at its maximum. 
     
     
         17 . The method of  claim 16 , wherein the value of the cutoff is −0.0575. 
     
     
         18 . The method of  claim 15 , wherein the gene marker set comprises at least three of the genes in SEQ ID NOs: 1-31. 
     
     
         19 . The method of  claim 15 , wherein the gene marker set comprises at least four of the genes in SEQ ID NOs: 1-31. 
     
     
         20 . The method of  claim 15 , wherein the gene marker set comprises the genes in SEQ ID NOs: 1-31.

Join the waitlist — get patent alerts

Track US2019367995A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.