US2006069512A1PendingUtilityA1

Gene discovery through comparisons of networks of structural and functional relationships among known genes and proteins

Assignee: RZHETSKY ANDREYPriority: Apr 15, 1999Filed: Aug 18, 2004Published: Mar 30, 2006
Est. expiryApr 15, 2019(expired)· nominal 20-yr term from priority
G16B 30/00G16B 40/00G16B 5/20G16B 50/00G16B 5/00G16B 20/00G16B 50/20G16B 30/10G16B 40/20G16B 10/00G16H 70/60Y02A90/10
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention relates to methods for identifying novel genes comprising: (i) generating one or more specialized databases containing information on gene/protein structure, function and/or regulatory interactions; and (ii) searching the specialized databases for homology or for a particular motif and thereby identifying a putative novel gene of interest. The invention may further comprise performing simulation and hypothesis testing to identify or confirm that the putative gene is a novel gene of interest. The present invention also relates to natural language processing and extraction of relational information associated with genes and proteins that are found in genomics journal articles. To enable access to information in textual form, the natural language processing system of the present invention provides a method for extracting and structuring information found in the literature in a form appropriate for subsequent applications.

Claims

exact text as granted — not AI-modified
1 . A method for identifying a novel nucleic acid molecule encoding a protein of interest comprising: 
 (i) selecting a specific protein from a first species involved in a regulatory network of interest;    (ii) identifying known proteins that act upstream and downstream in the regulatory network of interest with respect to the specific protein selected;    (iii) constructing the regulatory network of interest from the proteins identified in step (ii);    (iv) for each identified protein, select a domain or motif and search by homology for related proteins in a second species, wherein a related protein is defined as a protein having a homologous domain or motif;    (v) producing a regulatory network for the second species, wherein said regulatory network incorporates the identified related proteins;    (vi) comparing the regulatory network from the first species to the regulatory network of said second species;    (v) identifying a protein present in a regulatory network for one species but absent in the regulatory network of the other species; and    (vi) isolating a nucleic acid molecule encoding the protein identified in step (v) in the species in which it is missing.    
     
     
         2 . The method of  claim 1  wherein the nucleic acid molecule encodes human protein.  
     
     
         3 . The method of  claim 1  wherein the related proteins are orthologs.  
     
     
         4 . The method of  claim 1  wherein the regulatory pathway is involved in apoptosis.  
     
     
         5 . The method of  claim 1  wherein the specific protein from the first species is involved in tumor suppression.  
     
     
         6 . A method for identifying the affect of a gene knockout on a regulatory pathway comprising the following steps: 
 (i) identification of the shortest non-oriented pathway connecting two gene products;    (ii) assigning an initial sign value of “−” to the knockout since the knockout gene product is inactive;    (iii) moving along the shortest pathway between the two gene products multiplying the sign with the sign of the next gene product in the pathway, wherein “−” stands for inhibition, “+” stands for induction or activation, and “0” stands for the lack of interaction between two proteins in the specified direction; and    (iv) determining the final sign at the end of the pathway, wherein “−” indicates inhibition and “+” indicates induction or activation of the pathway.    
     
     
         7 . A method for identifying a novel nucleic acid molecule encoding a protein of interest comprising: 
 (i) selecting a gene of interest and searching a database for homologous sequences;    (ii) aligning the homologous sequences identified in step (i);    (iii) constructing a gene tree using the sequence alignment;    (iv) constructing a species tree;    (v) imputing the species tree and gene tree into an algorithm which integrates the species tree and the gene tree into a reconciled tree; and    (vi) identifying orthologous genes present in one species but missing in another.    
     
     
         8 . The method of  claim 7  wherein the following algorithm is used to integrate the species tree and the gene tree into a reconciled tree: 
 (i) computing the similarity σ(S gi ,S sj ) for each pair of interior nodes from trees T g  and T s ,    (ii) finding the maximum σ(S gi ,S sj );    (iii) saving S gi  as a new cluster of orthologs, save {S gi }-{S sj  } as a set of species that are likely to have gene of this kind (or lost it in evolution);    (iv) eliminating S gi  from T g ; T g :=T g \S gi ;    (v) repeating step (ii)-(iv) until T g  is non-empty.    
     
     
         9 . A method for identifying a novel gene comprising the following steps: 
 (i) defining a motif or domain composition of a gene of interest;    (ii) searching for sequences which correspond to nucleotide sequences in an expression sequence tag database or other cDNA databases using a program such as BLAST and retrieving the identified sequences;    (iii) searching additional databases for expressed sequence tags containing the domains and motifs characteristic for the gene of interest with Hidden Markov Model of domains and motifs identified in step (i);    (iv) identifying nucleotide sequences comprising the gene of interest.    
     
     
         10 . The method of  claim 9  further comprising using each identified expression sequence tag to search sequence databases for overlapping sequences for the purpose of assembling longer overlapping stretches of DNA.  
     
     
         11 - 21 . (canceled)  
     
     
         22 . A computer system for extracting information on biological entities from natural-language text data, comprising: 
 (i) means for parsing the natural-language text data; and    (ii) means for regularizing the parsed text data to form structured word terms.    
     
     
         23 . The system according to  claim 22 , further comprising means for preprocessing the data prior to parsing, with the preprocessing means comprising identifying biological entities.  
     
     
         24 . The system according to  claim 22 , further comprising means for referring to an additional parameter which is indicative of the degree to which subphrase parsing is to be carried out.  
     
     
         25 . The system according to  claim 22 , wherein said parsing means further comprises means for segmenting the text data by sentences.  
     
     
         26 . The system according to  claim 22 , wherein said parsing means further comprises: 
 means for segmenting the text data by sentences; and    means for segmenting each of the sentences at identified words or phrases.    
     
     
         27 . The system according to  claim 22 , wherein said parsing means further comprises: 
 means for segmenting the text data by sentences; and    means for segmenting each of the sentences at a prefix.    
     
     
         28 . The system according to  claim 22 , wherein said parsing means further comprises means for skipping undefined words.  
     
     
         29 . The system according to  claim 22 , wherein said parsing means further comprises: 
 means for identifying one or more binary actions and their relationships; and    means for identifying one or more arguments associated with the actions.    
     
     
         30 . The system according to  claim 22 , further comprising means for performing error recovery when parsing of the text data is unsuccessful.  
     
     
         31 . The system according to  claim 22 , wherein said error recovery means comprises: 
 means for segmenting the text data; and    means for analyzing the segmented text data to achieve at least a partial parsing of the unsuccessfully parsed text data.    
     
     
         32 . The system according to  claim 22 , wherein said tagging means comprises means for providing the structured data component in a Standard Generalized Markup Language (SGML) compatible format.

Join the waitlist — get patent alerts

Track US2006069512A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.