Methods and systems for linking data records from disparate databases
Abstract
In an illustrative embodiment, systems and methods for performing cascading matching of data records from disparate data sources comprise identifying matches using at least one uniquely identifying data field and at least one additional data field shared by a first data set and a second data set. Potential matches may be resolved through calculating differences between one or more shared data fields of a matched data record of the first data set and both a first matched record and a second matched record of the second data set, and determining a best match through analyzing the calculated differences. Unmatched records may be iteratively matched using a different uniquely identifying data field and/or different at least one additional data field(s).
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for matching data records obtained from disparate data stores, comprising:
a) identifying a plurality of overlapping data fields existing in both a first data set of data records and a second data set of data records; b) identifying, from the plurality of overlapping data fields, a plurality of identifier data fields containing uniquely identifying information and a plurality of remaining data fields of the plurality of overlapping data fields not containing uniquely identifying information; c) determining at least one identifier field of the plurality of identifier data fields and at least one remaining field of the plurality of remaining data fields for merging; d) merging, by processing circuitry using the at least one identifier field and the at least one remaining field, the first data set with the second data set to identify a plurality of data record matches; e) determining, by the processing circuitry, whether the plurality of data record matches comprises a plurality of potential data matches involving at least one same data record of one of the first data set and the second data set; f) responsive to determining the plurality of potential data matches,
calculating, by the processing circuitry, for each potential match of the plurality of potential data matches, differences between one or more shared data fields of the plurality of shared data fields, and
selecting, by the processing circuitry, a best match based upon the calculated differences;
g) identifying, by the processing circuitry, each matched data record of the first data set and the second data set as ineligible for further matching; and h) while at least one data record of the first data set and the second data set is not marked as ineligible for further matching and unused fields of the plurality of identifier data fields remain,
identifying, by the processing circuitry, at least one of i) a different one or more fields of the plurality of identifier data fields and ii) a different one or more fields of the plurality of remaining data fields for use in merging, by the processing circuitry, data records of the first data set and data records of the second data set not marked as ineligible for further matching, and
repeating steps (d) through (g).Join the waitlist — get patent alerts
Track US2021109953A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.