Topic models
Abstract
Machine learning techniques may be used to train computing devices to understand a variety of documents (e.g., text files, web pages, articles, spreadsheets, etc.). Machine learning techniques may be used to address the issue that computing devices may lack the human intellect used to understand such documents, such as their semantic meaning. Accordingly, a topic model may be trained by sequentially processing documents and/or their features (e.g., document author, geographical location of author, creation date, social network information of author, and/or document metadata). Additionally, as provided herein, the topic model may be used to predict probabilities that words, features, documents, and/or document corpora, for example, are indicative of particular topics.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for predicting a topic of a document, comprising:
processing a document representation of a document and features of the document using a topic model, the processing comprising:
determining a feature/topic prediction for a feature of the document, the feature/topic prediction specifying a probability of the feature being indicative of a first topic;
determining a word/topic prediction for a word within the document, the word/topic prediction specifying a probability of the word being indicative of a second topic; and
determining a document/topic prediction for the document based upon the feature/topic prediction and the word/topic prediction.
2 . The method of claim 1 , the document/topic prediction specifying a probability of the document being indicative of a third topic.
3 . The method of claim 1 , the features of the document comprising author designation.
4 . The method of claim 1 , the features of the document comprising geographical location.
5 . The method of claim 1 , the features of the document comprising at least one of creation date designation or creation time designation.
6 . The method of claim 1 , the features of the document comprising at least one of document metadata or source of the document.
7 . The method of claim 1 , the features of the document comprising document length.
8 . The method of claim 1 , the features of the document comprising social network membership.
9 . The method of claim 1 , the features of the document comprising previous search queries.
10 . The method of claim 1 , the features of the document comprising reader designation.
11 . The method of claim 1 , the features of the document comprising web browsing history.
12 . The method of claim 1 , the features of the document comprising document type.
13 . A system for training a topic model, comprising:
one or more processing units; and memory comprising instructions that when executed by at least one of the one or more processing units, perform a method comprising:
for a document within a document corpus:
receiving a document representation of the document and features of the document, the document representation comprising a frequency of word occurrences within the document;
processing the document representation and the features using a topic model, the processing comprising:
specifying a feature/topic parameter for a feature of the document, the feature/topic parameter specifying a probability of the feature being indicative of a first topic, the feature/topic parameter based upon a first uncertainty measure; and
specifying a document/word/topic parameter for a word within the document, the document/word/topic parameter specifying a probability of the word being indicative of a second topic, the document/word/topic parameter based upon a second uncertainty measure; and
training the topic model based upon the feature/topic parameter and the document/word/topic parameter.
14 . The system of claim 13 , the features of the document comprising at least one of author designation or reader designation.
15 . The system of claim 13 , the features of the document comprising geographical location.
16 . The system of claim 13 , the features of the document comprising at least one of creation date designation or creation time designation.
17 . The system of claim 13 , the features of the document comprising at least one of document metadata or source of the document.
18 . The system of claim 13 , the features of the document comprising at least one of document length or document type.
19 . The system of claim 13 , the features of the document comprising at least one of social network membership or web browsing history.
20 . A computer readable medium comprising instructions that when executed, perform a method for predicting a topic of a document, the method comprising:
processing a document representation of a document and features of the document using a topic model, the processing comprising:
determining a feature/topic prediction for a feature of the document, the feature/topic prediction specifying a probability of the feature being indicative of a first topic;
determining a word/topic prediction for a word within the document, the word/topic prediction specifying a probability of the word being indicative of a second topic; and
determining a document/topic prediction for the document based upon the feature/topic prediction and the word/topic prediction, the document/topic prediction specifying a probability of the document being indicative of a third topic.Join the waitlist — get patent alerts
Track US2014156571A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.