US2006277175A1PendingUtilityA1

Method and Apparatus for Focused Crawling

Assignee: JIANG DONGMINGPriority: Aug 18, 2000Filed: May 26, 2006Published: Dec 7, 2006
Est. expiryAug 18, 2020(expired)· nominal 20-yr term from priority
G06F 40/131Y10S707/99937G06F 16/951G06F 40/143
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention pertains to the field of computer software. More specifically, the present invention relates to dynamic discovery of documents or information through a focused crawler or search engine.

Claims

exact text as granted — not AI-modified
1 . A method of ranking documents, comprising: 
 accessing a first plurality of documents from a database of a plurality of received documents, the first plurality of documents to be ranked;    generating a graph of the first plurality of documents;    expanding the graph with a second plurality of one or more documents from the database, such that a third plurality includes a union of the first plurality of documents and the second plurality of documents, and the third plurality of documents is smaller than the plurality of received documents, the second plurality including one or more of: 1) one or more documents connected within a first specified number of links in a forward direction from one or more documents of the first plurality of documents, the forward direction being forward from the first plurality of documents, and 2) one or more documents connected within a second specified number of links in a backward direction from one or more documents of the first plurality of documents, the backward direction being backward from the first plurality of documents;    assigning weights to one or more nodes of the graph;    finding an assignment of weights to one or more nodes of the graph, by propagating weights through the graph, the assignment of weight to a node based at least in part on calculating a weighted sum of weights propagated from neighboring nodes; and    generating a ranked list of at least the first plurality of documents, the ranked list at least partly generated from the graph.    
   
   
       2 . The method of  claim 1 , further comprising: 
 shrinking the graph by removing one or more nodes of the graph.    
   
   
       3 . The method of  claim 1 , further comprising: 
 shrinking the graph by combining one or more sets of one or more nodes of the graph.    
   
   
       4 . The method of  claim 3 , wherein the combining is based on common characteristics of the nodes or relationships between the nodes.  
   
   
       5 . The method of  claim 1 , wherein the propagating weights through the graph occurs up to a limited node distance.  
   
   
       6 . The method of  claim 1 , wherein weights assigned to a node include at least one of relevance of the document to a query input and importance of the document independent of the query input.  
   
   
       7 . A method of finding starting points, comprising: 
 accessing a plurality of sample documents, the sample documents including links to other documents;    generating a graph representation of the plurality of sample documents and a plurality of referring documents, each of the referring documents referring to at least one of the plurality of sample documents, such that nodes of the graph represent the sample documents and the referring documents, and edges of the graph represent links linking the sample documents and the referring documents; and    finding starting point documents, at least partly by performing a link structure analysis on the graph.    
   
   
       8 . The method of  claim 7 , wherein the link structure analysis includes: 
 assigning weights at least to nodes representing sample documents; and    propagating weights through the graph in reverse direction, to nodes representing referring pages; and    assigning weights to nodes of the graph, by calculating a weighted sum of weights propagated from neighboring nodes.    
   
   
       9 . The method of  claim 7 , wherein at least one of the plurality of referring documents refers to at least one of the plurality of sample documents directly.  
   
   
       10 . The method of  claim 7 , wherein at least one of the plurality of referring documents refers to at least one of the plurality of sample documents indirectly through one or more documents.  
   
   
       11 . The method of  claim 7 , wherein finding starting point documents further includes one or more of: 1) evaluating relevance of documents using logical expressions of keywords and phrases and 2) evaluating relevance of documents using content located in a specified part of a format.  
   
   
       12 . The method of  claim 7 , wherein relevance includes importance.  
   
   
       13 . The method of  claim 7 , wherein the starting point documents provide starting points for at least one crawl.  
   
   
       14 . The method of  claim 7 , wherein evaluating relevance of documents includes evaluating relevance of at least a first document and a second document, the second document referring to the first document.

Join the waitlist — get patent alerts

Track US2006277175A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.