US2012265521A1PendingUtilityA1

Methods and systems relating to information extraction

Assignee: MILLER SCOTTPriority: May 5, 2005Filed: Apr 16, 2012Published: Oct 18, 2012
Est. expiryMay 5, 2025(expired)· nominal 20-yr term from priority
G06F 40/211G06F 40/295
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The invention relates to information extraction systems having discriminative models which utilize hierarchical cluster trees and active learning to enhance training.

Claims

exact text as granted — not AI-modified
1 . An information extraction system comprising:
 a discriminative information extraction model implemented as computer readable instructions executable on one or more general or special purpose computers, the discriminative information extraction model comprising a classification set, a feature set, and a cluster tree; and   a training module implemented as computer readable instructions executable on the one or more general or special purpose computers, the training module comprising an active learning user interface and a data input.   
     
     
         2 . The system of  claim 1  wherein the classification set comprises:
 a plurality of classifications to which words in a sentence can be assigned; and 
 two non-name classifications referred to as none and null. 
 
     
     
         3 . The system of  claim 1  wherein the none classification refers to words in a sentence that are not part of a named entity. 
     
     
         4 . The system of  claim 1  where the null classification refers Null to a beginning of a sentence, before a first word, or an end of a sentence, after a last word. 
     
     
         5 . The system of  claim 1  wherein the feature set comprises a plurality of variables related to each classification and tag in the discriminative information extraction model. 
     
     
         6 . The system of claim  23  wherein the feature set comprises:
 word context features that relate to a use of specific words and related classifications: and 
 cluster context features that relate to the use of groups of similarly used words as defined by the cluster tree. 
 
     
     
         7 . The system of  claim 1  wherein the cluster tree groups words into hierarchical groups that have decreasingly similar usage statistics as the groups increase in size. 
     
     
         8 . The system of claim  34  wherein the discriminative information extraction model further comprises a data store including a training set and a feature weight table. 
     
     
         9 . The system of  claim 8  wherein the training set comprises a copy of all annotated text on which the discriminative information extraction model has been trained. 
     
     
         10 . The system of  claim 8  wherein the feature weight table stores weights corresponding to each of the features for each classification and tag in discriminative information extraction model. 
     
     
         11 . The system of  claim 10  wherein the discriminative information extraction model further comprises an evaluation module for determining a proper classification of a given word based on data stored in the feature weight table and in the cluster tree. 
     
     
         12 . The system of  claim 11  wherein the evaluation module determines a set of all possible tag sequences for the string of words and the set into a lattice structure, evaluates each path through the lattice, and selects a best path according to a Viterbi algorithm. 
     
     
         13 . The system of  claim 1  wherein the active learning user interface provides a graphical user interface (GUI) for labeling each word in a selected sentence.

Join the waitlist — get patent alerts

Track US2012265521A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.