US2011125747A1PendingUtilityA1

Data classification based on point-of-view dependency

Assignee: BIZ360 INCPriority: Aug 28, 2003Filed: Jun 24, 2010Published: May 26, 2011
Est. expiryAug 28, 2023(expired)· nominal 20-yr term from priority
G06F 16/353
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Data classification is used to classified input items by associating the input items with one or more classes from a set of one or more classes in a data classification system, including identifying relevant features in an input item to form a feature vector for the input item, receiving at the data classification system an indication of a point-of-view, adjusting the feature vector according to the point-of-view indication or modifying a pattern discriminator (e.g., trainer and classifier) to inline-process feature vectors depending on the provided point-of-view (e.g., SVM custom kernels), and classifying the input item into the set of classes according to the point-of-view. The point-of-view data can be introduced either as a pre-process step prior to passing it off to the pattern discrimination algorithm, or can be incorporated directly into the pattern discrimination algorithm if applicable. The pattern discrimination algorithms can detect arbitrary patterns given a similarly prepared dataset during both training and subsequent classification of unclassified documents.

Claims

exact text as granted — not AI-modified
1 . A method of classification, wherein an input item is classified by associating the input item with one or more classes from a set of classes in a data classification system, said method comprising the steps of:
 receiving the input item to be classified;   identifying relevant features in the input item to form a feature vector for the input item;   receiving an indication of a point of view at the data classification system;   adjusting the feature vector or modifying a pattern discriminator according to the point-of-view indication; and   classifying the input item into the set of classes according to the point-of-view.   
     
     
         2 . The method of  claim 1 , wherein the step of adjusting the feature vector comprises generating custom features. 
     
     
         3 . The method of  claim 1 , wherein the step of adjusting the feature vector comprises selecting a subset of features. 
     
     
         4 . The method of  claim 1 , wherein the step of modifying a pattern discrimination algorithm comprises generating a custom kernel. 
     
     
         5 . The method of  claim 1 , wherein the step of adjusting the feature vector comprises weighting features. 
     
     
         6 . The method of  claim 5 , wherein weighting features uses proximity weighting. 
     
     
         7 . The method of  claim 6 , wherein proximity weighting calculates weight of a feature as the maximum of 0.95 raised to the power of FSP and 0.80 raised to the power of BSP, wherein FSP is the number of sentences going forward from a nearest alias to the feature and BSP is the number of sentences going backward from a nearest alias to the feature, wherein an alias is a representation of a point-of-view. 
     
     
         8 . The method of  claim 1 , wherein the input item is selected from the group consisting of a word processing document, an ASCII file, an XML file, a UTF-8 file, a collection of documents with some structural organization, an image, a text, a combination of images and text, media, spreadsheet data, a collection of bytes, an organization of data and a data stream. 
     
     
         9 . A data classification system comprising at least one input item, at least one feature vector, and at least one data classifier defined by point-of-view dependency, wherein the data classification system is configured to perform one or more of feature generation, feature selection, feature weighting, and custom kernel generation in order to rate and classify the input item. 
     
     
         10 . The data classification system of  claim 9 , wherein the input item is selected from the group consisting of a word processing document, an ASCII file, an XML file, a UTF-8 file, a collection of documents with some structural organization, an image, a text, a combination of images and text, media, spreadsheet data, a collection of bytes, an organization of data and a data stream. 
     
     
         11 . The data classification system of  claim 9 , wherein the data classifier classifies one or more data sets based upon patterns observed during a training process with one or more training data sets.

Join the waitlist — get patent alerts

Track US2011125747A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.