Data classification based on point-of-view dependency
Abstract
Data classification is used to classified input items by associating the input items with one or more classes from a set of one or more classes in a data classification system, including identifying relevant features in an input item to form a feature vector for the input item, receiving at the data classification system an indication of a point-of-view, adjusting the feature vector according to the point-of-view indication or modifying a pattern discriminator (e.g., trainer and classifier) to inline-process feature vectors depending on the provided point-of-view (e.g., SVM custom kernels), and classifying the input item into the set of classes according to the point-of-view. The point-of-view data can be introduced either as a pre-process step prior to passing it off to the pattern discrimination algorithm, or can be incorporated directly into the pattern discrimination algorithm if applicable. The pattern discrimination algorithms can detect arbitrary patterns given a similarly prepared dataset during both training and subsequent classification of unclassified documents.
Claims
exact text as granted — not AI-modified1 . A method of classification, wherein an input item is classified by associating the input item with one or more classes from a set of classes in a data classification system, said method comprising the steps of:
receiving the input item to be classified; identifying relevant features in the input item to form a feature vector for the input item; receiving an indication of a point of view at the data classification system; adjusting the feature vector or modifying a pattern discriminator according to the point-of-view indication; and classifying the input item into the set of classes according to the point-of-view.
2 . The method of claim 1 , wherein the step of adjusting the feature vector comprises generating custom features.
3 . The method of claim 1 , wherein the step of adjusting the feature vector comprises selecting a subset of features.
4 . The method of claim 1 , wherein the step of modifying a pattern discrimination algorithm comprises generating a custom kernel.
5 . The method of claim 1 , wherein the step of adjusting the feature vector comprises weighting features.
6 . The method of claim 5 , wherein weighting features uses proximity weighting.
7 . The method of claim 6 , wherein proximity weighting calculates weight of a feature as the maximum of 0.95 raised to the power of FSP and 0.80 raised to the power of BSP, wherein FSP is the number of sentences going forward from a nearest alias to the feature and BSP is the number of sentences going backward from a nearest alias to the feature, wherein an alias is a representation of a point-of-view.
8 . The method of claim 1 , wherein the input item is selected from the group consisting of a word processing document, an ASCII file, an XML file, a UTF-8 file, a collection of documents with some structural organization, an image, a text, a combination of images and text, media, spreadsheet data, a collection of bytes, an organization of data and a data stream.
9 . A data classification system comprising at least one input item, at least one feature vector, and at least one data classifier defined by point-of-view dependency, wherein the data classification system is configured to perform one or more of feature generation, feature selection, feature weighting, and custom kernel generation in order to rate and classify the input item.
10 . The data classification system of claim 9 , wherein the input item is selected from the group consisting of a word processing document, an ASCII file, an XML file, a UTF-8 file, a collection of documents with some structural organization, an image, a text, a combination of images and text, media, spreadsheet data, a collection of bytes, an organization of data and a data stream.
11 . The data classification system of claim 9 , wherein the data classifier classifies one or more data sets based upon patterns observed during a training process with one or more training data sets.Join the waitlist — get patent alerts
Track US2011125747A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.