US2013097168A1PendingUtilityA1

Method to identify common structures in formatted text documents

Assignee: CHANG YUAN-CHIPriority: Dec 9, 2009Filed: Apr 5, 2012Published: Apr 18, 2013
Est. expiryDec 9, 2029(~3.4 yrs left)· nominal 20-yr term from priority
G06F 16/93G06F 16/353G06F 40/186G06F 17/30011
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A computer implemented method, computer program product and data processing system, for identifying common structures shared across a plurality of formatted text documents. The common structure is presented as a sequence of landmarks, each of which has a starting and ending marker to describe the borders of text. The common structure is identified by counting the occurrences of repeating text segments across documents. Frequently co-occurred adjacent segments become candidates for markers of landmarks. In addition, styling information of textual content within a landmark is extracted and mapped to rules. The rules are used to merge and summarize content from multiple documents, which gives an advantage over current practice of content concatenation.

Claims

exact text as granted — not AI-modified
Having thus described our invention, what we claim as new and desire to secure by Letters Patent is as follows: 
     
         1 . A computerized method to discover hidden structures in documents stored in a repository or document collection, said method comprising:
 retrieving documents from said repository, each retrieved document having one or more previously-identified markers, each said marker potentially serving as a basis for a template entry;   clustering, as executed by a processor on a computer, said retrieved documents into a plurality of clusters as based on a preset threshold of a number of markers that are shared by said retrieved documents, each said cluster representing a potential document template;   and selecting from said plurality of clusters, clusters that exceed a minimal cluster size, said selected clusters being output as comprising distinct document templates represented by the documents in said repository.   
     
     
         2 . The method of  claim 1 , wherein said clusters are a selected based on one of:
 an absolute number;   a fraction of the retrieved documents; and   a fraction of a total number of documents in said repository.   
     
     
         3 . The method of  claim 2 , further comprising counting and reporting on the distinct document templates. 
     
     
         4 . The method of  claim 2 , further comprising preliminarily determining said markers on said documents. 
     
     
         5 . The method of  claim 2 , wherein weights are assigned to said shared markers used for said clustering. 
     
     
         6 . The method of  claim 2 , wherein all of said documents in said repository are retrieved for said clustering. 
     
     
         7 . The method of claim wherein only a portion of said documents in said repository are retrieved for said clustering, as representative of said repository. 
     
     
         8 . The method of  claim 7 , wherein said portion of documents retrieved are selected randomly. 
     
     
         9 . The method of  claim 2 , further comprising:
 retrieving one or more additional documents from said repository;   for each additional retrieved document, extracting a content from said retrieved document; and   using said extracted content to verify one or more of said distinct document templates.   
     
     
         10 . The method of  claim 1 , as comprising a set of machine readable instructions tangibly embodied in a tangible machine readable storage medium.

Join the waitlist — get patent alerts

Track US2013097168A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.