US2014156571A1PendingUtilityA1

Topic models

Assignee: MICROSOFT CORPPriority: Oct 26, 2010Filed: Feb 4, 2014Published: Jun 5, 2014
Est. expiryOct 26, 2030(~4.2 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 20/00G06N 7/005G06N 99/005
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Machine learning techniques may be used to train computing devices to understand a variety of documents (e.g., text files, web pages, articles, spreadsheets, etc.). Machine learning techniques may be used to address the issue that computing devices may lack the human intellect used to understand such documents, such as their semantic meaning. Accordingly, a topic model may be trained by sequentially processing documents and/or their features (e.g., document author, geographical location of author, creation date, social network information of author, and/or document metadata). Additionally, as provided herein, the topic model may be used to predict probabilities that words, features, documents, and/or document corpora, for example, are indicative of particular topics.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for predicting a topic of a document, comprising:
 processing a document representation of a document and features of the document using a topic model, the processing comprising:
 determining a feature/topic prediction for a feature of the document, the feature/topic prediction specifying a probability of the feature being indicative of a first topic; 
 determining a word/topic prediction for a word within the document, the word/topic prediction specifying a probability of the word being indicative of a second topic; and 
 determining a document/topic prediction for the document based upon the feature/topic prediction and the word/topic prediction. 
   
     
     
         2 . The method of  claim 1 , the document/topic prediction specifying a probability of the document being indicative of a third topic. 
     
     
         3 . The method of  claim 1 , the features of the document comprising author designation. 
     
     
         4 . The method of  claim 1 , the features of the document comprising geographical location. 
     
     
         5 . The method of  claim 1 , the features of the document comprising at least one of creation date designation or creation time designation. 
     
     
         6 . The method of  claim 1 , the features of the document comprising at least one of document metadata or source of the document. 
     
     
         7 . The method of  claim 1 , the features of the document comprising document length. 
     
     
         8 . The method of  claim 1 , the features of the document comprising social network membership. 
     
     
         9 . The method of  claim 1 , the features of the document comprising previous search queries. 
     
     
         10 . The method of  claim 1 , the features of the document comprising reader designation. 
     
     
         11 . The method of  claim 1 , the features of the document comprising web browsing history. 
     
     
         12 . The method of  claim 1 , the features of the document comprising document type. 
     
     
         13 . A system for training a topic model, comprising:
 one or more processing units; and   memory comprising instructions that when executed by at least one of the one or more processing units, perform a method comprising:
 for a document within a document corpus:
 receiving a document representation of the document and features of the document, the document representation comprising a frequency of word occurrences within the document; 
 processing the document representation and the features using a topic model, the processing comprising:
 specifying a feature/topic parameter for a feature of the document, the feature/topic parameter specifying a probability of the feature being indicative of a first topic, the feature/topic parameter based upon a first uncertainty measure; and 
 specifying a document/word/topic parameter for a word within the document, the document/word/topic parameter specifying a probability of the word being indicative of a second topic, the document/word/topic parameter based upon a second uncertainty measure; and 
 
 training the topic model based upon the feature/topic parameter and the document/word/topic parameter. 
 
   
     
     
         14 . The system of  claim 13 , the features of the document comprising at least one of author designation or reader designation. 
     
     
         15 . The system of  claim 13 , the features of the document comprising geographical location. 
     
     
         16 . The system of  claim 13 , the features of the document comprising at least one of creation date designation or creation time designation. 
     
     
         17 . The system of  claim 13 , the features of the document comprising at least one of document metadata or source of the document. 
     
     
         18 . The system of  claim 13 , the features of the document comprising at least one of document length or document type. 
     
     
         19 . The system of  claim 13 , the features of the document comprising at least one of social network membership or web browsing history. 
     
     
         20 . A computer readable medium comprising instructions that when executed, perform a method for predicting a topic of a document, the method comprising:
 processing a document representation of a document and features of the document using a topic model, the processing comprising:
 determining a feature/topic prediction for a feature of the document, the feature/topic prediction specifying a probability of the feature being indicative of a first topic; 
 determining a word/topic prediction for a word within the document, the word/topic prediction specifying a probability of the word being indicative of a second topic; and 
 determining a document/topic prediction for the document based upon the feature/topic prediction and the word/topic prediction, the document/topic prediction specifying a probability of the document being indicative of a third topic.

Join the waitlist — get patent alerts

Track US2014156571A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.