Method of reducing redundancy between two or more datasets
Abstract
A method for reducing redundancy between two or more datasets of potentially very large size. The method improves upon current technology by oversubscribing the data structure that represents a digest of data blocks and using positional information about matching data so that very large datasets can be analyzed and the redundancies removed by, having found a match on digest, expands the match in both directions in order to detect and eliminate large runs of data by replace duplicate runs with references to common data. The method is particularly useful for capturing the states of images of a hard disk. The method permits several files to have their redundancy removed and the files to later be reconstituted. The method is appropriate for use on a WORM device. The method can also make use of L2 cache to improve performance.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of reducing redundancy between two or more data sets comprising:
generating a plurality of first hash codes for a plurality of data blocks associated with one or more reference files, wherein the first hash codes are generated with a first hash algorithm executing in one or more computer processors; storing one or more of the plurality of first hash codes in one or more hash entries in a hash table; using the first hash algorithm to compute at least a first hash code for a current data block associated with a current file; comparing the first hash code associated with the current data block with the one or more first hash codes stored in the one or more hash entries in the hash table; when the first hash code of the current data block matches at least one of the first hash codes in the one or more hash entries in the hash table, generating a second hash code for the current data block, wherein the second hash code is generated with a second hash algorithm that is computationally more expensive than the first hash algorithm; and when the second hash code for the current data block matches a second hash code associated with the first hash entry, comparing the data in at least one of preceding and succeeding data blocks of the current and reference files to identify a matching run of data in the current and reference files.Join the waitlist — get patent alerts
Track US2019340165A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.