US2006179050A1PendingUtilityA1

Probabilistic model for record linkage

Individually held — no corporate assignee on recordPriority: Oct 22, 2004Filed: Oct 21, 2005Published: Aug 10, 2006
Est. expiryOct 22, 2024(expired)· nominal 20-yr term from priority
G06F 16/3346
37
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for probabilistic record linkage includes providing a record pair comprising a plurality of fields, providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses, and comparing the record pair to determine a record difference. The method includes determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden, determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes, and outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method for probabilistic record linkage comprising: 
 providing a record pair comprising a plurality of fields;    providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses;    comparing the record pair to determine a record difference;    determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden;    determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes; and    outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.    
   
   
       2 . The computer-implemented method of  claim 1 , wherein comparing the record pair comprises comparing record values of the record pair field-wise or across fields.  
   
   
       3 . The computer-implemented method of  claim 1 , wherein determining the probability of a status for each of the plurality of attributes comprises: 
 providing a predefined error rate of data entering in a field;    determining a distance metric between field values; and    determining a probability of making i errors when entering m characters with the predefined error rate.    
   
   
       4 . The computer-implemented method of  claim 1 , wherein each among a plurality of scenarios is characterized by a probability model on patterns of attribute statuses for example Bayesian net, conditional probabilities of attribute status given scenarios.  
   
   
       5 . The computer-implemented method of  claim 1 , 
 wherein the probability of duplication is compared to a threshold, wherein the threshold corresponds to a significant probability of duplication.    
   
   
       6 . The computer-implemented method of  claim 1 , further comprising: 
 providing a graphical user interface; and    displaying at least one of a scenario probability, a most probable scenario, a probability of duplication, and/or a probability that an entity is intended by an input search criteria.    
   
   
       7 . The computer-implemented method of  claim 1 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising: 
 determining for each database record the probability of duplication or non-duplication as a probability that the record is the target of the search criteria; and    displaying in a graphical user interface the database records and a corresponding probability.    
   
   
       8 . The computer-implemented method of  claim 1 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising: 
 determining for each database record the probability of duplication or non-duplication as a confidence score corresponding to the search criteria; and    displaying in a graphical user interface each database records and a corresponding confidence score.    
   
   
       9 . A computer-implemented method comprising: 
 receiving a record pair; and    outputting a probability of duplication between the record pair from an observation of field values of the record pair and noisy characteristics of the record pair.    
   
   
       10 . The computer-implemented method of  claim 9 , wherein the observation of field values is one of an edit distance, a soundex distance, a numerical distance, or a date distance between a pair of fields corresponding to the record pair, respectively.  
   
   
       11 . The computer-implemented method of  claim 9 , further comprising modeling the noisy characteristics of the record pair comprising: 
 determining a probability of a difference between attribute values corresponding to the fields; and    determining a probability of an error in the field values.    
   
   
       12 . A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for probabilistic record linkage, the method steps comprising: 
 providing a record pair comprising a plurality of fields;    providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses;    comparing the record pair to determine a record difference;    determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden;    determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes; and    outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.    
   
   
       13 . The method of  claim 12 , wherein comparing the record pair comprises comparing record values of the record pair field-wise or across fields.  
   
   
       14 . The method of  claim 12 , wherein determining the probability of a status for each of the plurality of attributes comprises: 
 providing a predefined error rate of data entering in a field;    determining a distance metric between field values; and    determining a probability of making i errors when entering m characters with the predefined error rate.    
   
   
       15 . The method of  claim 12 , wherein each among a plurality of scenarios is characterized by a probability model on patterns of attribute statuses for example Bayesian net, conditional probabilities of attribute status given scenarios.  
   
   
       16 . The method of  claim 12 , wherein the probability of duplication is compared to a threshold, wherein the threshold corresponds to a significant probability of duplication.  
   
   
       17 . The method of  claim 12 , further comprising: 
 providing a graphical user interface; and    displaying at least one of a scenario probability, a most probable scenario, a probability of duplication, and/or a probability that an entity is intended by an input search criteria.    
   
   
       18 . The method of  claim 12 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising: 
 determining for each database record the probability of duplication or non-duplication as a probability that the record is the target of the search criteria; and    displaying in a graphical user interface the database records and a corresponding probability.    
   
   
       19 . The method of  claim 11 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising: 
 determining for each database record the probability of duplication or non-duplication as a confidence score corresponding to the search criteria; and    displaying in a graphical user interface each database records and a corresponding confidence score.

Join the waitlist — get patent alerts

Track US2006179050A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.