US2012265521A1PendingUtilityA1
Methods and systems relating to information extraction
Est. expiryMay 5, 2025(expired)· nominal 20-yr term from priority
Inventors:Scott Michael Miller
G06F 40/211G06F 40/295
51
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The invention relates to information extraction systems having discriminative models which utilize hierarchical cluster trees and active learning to enhance training.
Claims
exact text as granted — not AI-modified1 . An information extraction system comprising:
a discriminative information extraction model implemented as computer readable instructions executable on one or more general or special purpose computers, the discriminative information extraction model comprising a classification set, a feature set, and a cluster tree; and a training module implemented as computer readable instructions executable on the one or more general or special purpose computers, the training module comprising an active learning user interface and a data input.
2 . The system of claim 1 wherein the classification set comprises:
a plurality of classifications to which words in a sentence can be assigned; and
two non-name classifications referred to as none and null.
3 . The system of claim 1 wherein the none classification refers to words in a sentence that are not part of a named entity.
4 . The system of claim 1 where the null classification refers Null to a beginning of a sentence, before a first word, or an end of a sentence, after a last word.
5 . The system of claim 1 wherein the feature set comprises a plurality of variables related to each classification and tag in the discriminative information extraction model.
6 . The system of claim 23 wherein the feature set comprises:
word context features that relate to a use of specific words and related classifications: and
cluster context features that relate to the use of groups of similarly used words as defined by the cluster tree.
7 . The system of claim 1 wherein the cluster tree groups words into hierarchical groups that have decreasingly similar usage statistics as the groups increase in size.
8 . The system of claim 34 wherein the discriminative information extraction model further comprises a data store including a training set and a feature weight table.
9 . The system of claim 8 wherein the training set comprises a copy of all annotated text on which the discriminative information extraction model has been trained.
10 . The system of claim 8 wherein the feature weight table stores weights corresponding to each of the features for each classification and tag in discriminative information extraction model.
11 . The system of claim 10 wherein the discriminative information extraction model further comprises an evaluation module for determining a proper classification of a given word based on data stored in the feature weight table and in the cluster tree.
12 . The system of claim 11 wherein the evaluation module determines a set of all possible tag sequences for the string of words and the set into a lattice structure, evaluates each path through the lattice, and selects a best path according to a Viterbi algorithm.
13 . The system of claim 1 wherein the active learning user interface provides a graphical user interface (GUI) for labeling each word in a selected sentence.Join the waitlist — get patent alerts
Track US2012265521A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.