US2016364468A1PendingUtilityA1

Database index for constructing large scale data level of details

Assignee: IBMPriority: Jun 10, 2015Filed: Jan 4, 2016Published: Dec 15, 2016
Est. expiryJun 10, 2035(~8.9 yrs left)· nominal 20-yr term from priority
G06F 16/285G06F 16/2228G06F 16/2246G06F 17/30598G06F 17/30321
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An index for large databases is disclosed. Data is grouped into clusters and the clusters are grouped into levels of detail. Analysis results are determined based on progressive data sampling. Sampling is conducted based on the level of detail required and/or the resources (time or computing resources) that are available. Larger, more concentrated clusters, at higher levels of detail, are sampled more sparsely. Smaller, more diffuse clusters, at lower levels of detail, are sampled more intensively. Analysis results, including outlier data, include proportional representation from the whole database up to the level of detail required. Results are quickly determined with specified degree of accuracy, based on initial sampling, and are refined with subsequent sampling.

Claims

exact text as granted — not AI-modified
1 . A method comprising:
 receiving a plurality of record identifiers with each record identifier uniquely identifying a record in a database;   performing cluster analysis on the records corresponding to the plurality of record identifiers to yield a plurality of clusters, with each cluster including at least one record;   constructing a database index data structure where each record identifier is represented as a leaf node, each cluster is represented as a non-leaf node, and each leaf node is related to at least one non-leaf node based upon which record identifiers belong to which clusters; and   selecting a sample of records to be searched using the database index data structure, with the selection of records in the sample based, at least in part, on: (i) the density information of the non-leaf nodes of the database index data structure, (ii) the cardinality information of the non-leaf nodes of the database index data structure, (iii) the level information of the non-leaf nodes of the database index data structure, and (iv) the range information of the non-leaf nodes of the database index data structure;   wherein at least some of the non-leaf nodes include:
 density information corresponding to a density of the cluster corresponding to the non-leaf node, 
 cardinality information corresponding to a cardinality of the cluster corresponding to the non-leaf node, 
 level information corresponding to a level of detail of the cluster corresponding to the non-leaf node, and 
 range information corresponding to a data range of the cluster corresponding to the non-leaf node.

Join the waitlist — get patent alerts

Track US2016364468A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.