Probabilistic model for record linkage
Abstract
A method for probabilistic record linkage includes providing a record pair comprising a plurality of fields, providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses, and comparing the record pair to determine a record difference. The method includes determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden, determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes, and outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for probabilistic record linkage comprising:
providing a record pair comprising a plurality of fields; providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses; comparing the record pair to determine a record difference; determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden; determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes; and outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.
2 . The computer-implemented method of claim 1 , wherein comparing the record pair comprises comparing record values of the record pair field-wise or across fields.
3 . The computer-implemented method of claim 1 , wherein determining the probability of a status for each of the plurality of attributes comprises:
providing a predefined error rate of data entering in a field; determining a distance metric between field values; and determining a probability of making i errors when entering m characters with the predefined error rate.
4 . The computer-implemented method of claim 1 , wherein each among a plurality of scenarios is characterized by a probability model on patterns of attribute statuses for example Bayesian net, conditional probabilities of attribute status given scenarios.
5 . The computer-implemented method of claim 1 ,
wherein the probability of duplication is compared to a threshold, wherein the threshold corresponds to a significant probability of duplication.
6 . The computer-implemented method of claim 1 , further comprising:
providing a graphical user interface; and displaying at least one of a scenario probability, a most probable scenario, a probability of duplication, and/or a probability that an entity is intended by an input search criteria.
7 . The computer-implemented method of claim 1 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising:
determining for each database record the probability of duplication or non-duplication as a probability that the record is the target of the search criteria; and displaying in a graphical user interface the database records and a corresponding probability.
8 . The computer-implemented method of claim 1 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising:
determining for each database record the probability of duplication or non-duplication as a confidence score corresponding to the search criteria; and displaying in a graphical user interface each database records and a corresponding confidence score.
9 . A computer-implemented method comprising:
receiving a record pair; and outputting a probability of duplication between the record pair from an observation of field values of the record pair and noisy characteristics of the record pair.
10 . The computer-implemented method of claim 9 , wherein the observation of field values is one of an edit distance, a soundex distance, a numerical distance, or a date distance between a pair of fields corresponding to the record pair, respectively.
11 . The computer-implemented method of claim 9 , further comprising modeling the noisy characteristics of the record pair comprising:
determining a probability of a difference between attribute values corresponding to the fields; and determining a probability of an error in the field values.
12 . A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for probabilistic record linkage, the method steps comprising:
providing a record pair comprising a plurality of fields; providing a plurality of scenarios, each scenario relating to a distribution of patterns among a plurality of attribute statuses; comparing the record pair to determine a record difference; determining a probability of a status for each of a plurality of attributes based on the distance metric of the plurality of fields, wherein each field corresponds to a respective attribute, wherein the field is observable and the attribute is hidden; determining a probability of each scenario based on the probability of the status for each attribute and the Bayesian net representing the probabilistic model on the relationship between scenarios and attributes; and outputting a probability of duplication or non-duplication of the record pair determined from the probabilities of the plurality of scenarios.
13 . The method of claim 12 , wherein comparing the record pair comprises comparing record values of the record pair field-wise or across fields.
14 . The method of claim 12 , wherein determining the probability of a status for each of the plurality of attributes comprises:
providing a predefined error rate of data entering in a field; determining a distance metric between field values; and determining a probability of making i errors when entering m characters with the predefined error rate.
15 . The method of claim 12 , wherein each among a plurality of scenarios is characterized by a probability model on patterns of attribute statuses for example Bayesian net, conditional probabilities of attribute status given scenarios.
16 . The method of claim 12 , wherein the probability of duplication is compared to a threshold, wherein the threshold corresponds to a significant probability of duplication.
17 . The method of claim 12 , further comprising:
providing a graphical user interface; and displaying at least one of a scenario probability, a most probable scenario, a probability of duplication, and/or a probability that an entity is intended by an input search criteria.
18 . The method of claim 12 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising:
determining for each database record the probability of duplication or non-duplication as a probability that the record is the target of the search criteria; and displaying in a graphical user interface the database records and a corresponding probability.
19 . The method of claim 11 , wherein the record pair is a search criteria for determining a target and a plurality of database records, the method further comprising:
determining for each database record the probability of duplication or non-duplication as a confidence score corresponding to the search criteria; and displaying in a graphical user interface each database records and a corresponding confidence score.Join the waitlist — get patent alerts
Track US2006179050A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.