US2021279215A1PendingUtilityA1

Systems and methods for providing data quality management

Assignee: CAPITAL ONE SERVICES LLCPriority: Dec 19, 2016Filed: May 21, 2021Published: Sep 9, 2021
Est. expiryDec 19, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06F 16/2365G06F 16/122G06F 16/48G06N 20/00G06N 5/025G06F 16/215
60
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system for providing data quality management may include a processor configured to execute instructions to: extract a plurality of first data elements from a data source; generate a data profile based on the first data elements; automatically create a first set of rules based on the first data elements and the data profile, the first set of rules assessing data quality according to a threshold; generate a second set of rules based on the first data elements and the first set of rules; extract a plurality of second data elements; assess the second data elements based on a comparison of the second data elements to the second set of rules; detect defects based on the comparison; analyze data quality according to the detected defects; and transmit signals representing the data quality analysis to a client device for display to a user.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system for providing data quality management, the system comprising:
 at least one memory storing instructions; and   at least one processor connected to a network and executing the instructions to perform operations comprising:
 obtaining a set of rules for assessing data quality of a set of data elements; 
 extracting a plurality of data elements from a data source; 
 assessing the data elements based on a comparison of the data elements to the set of rules; 
 detecting a plurality of defects based on the comparison; 
 determining, using a decision tree algorithm, data quality according to the detected defects, the data quality comprising a pocket of defect concentration; and 
 transmitting, to a client device, instructions, to display a representation of the determined pocket of defect concentration in a user interface of the client device. 
   
     
     
         2 . The system of  claim 1 , wherein the operations further comprise:
 clustering the data elements into multiple segments based on the defects.   
     
     
         3 . The system of  claim 1 , wherein assessing the data elements includes:
 generating a first set of rules based on a property of the set of data elements;   generating a second set of rules based on the set of data elements and the first set of rules; and   assessing the data elements based on the second set of rules.   
     
     
         4 . The system of  claim 3 , wherein the first set of rules includes a null rule, a range rule, a uniqueness rule, a valid value rule, a format rule, a conditional rule, or a consistency rule, and wherein the second set of rules includes a framework criteria comprising a support value, a confidence value, or a lift ratio. 
     
     
         5 . A method for providing data quality management, the method comprising:
 obtaining a set of rules for assessing data quality of a set of data elements;   extracting a plurality of data elements from a data source;   assessing the data elements based on a comparison of the data elements to the set of rules;   detecting a plurality of defects based on the comparison;   determining data quality according to the detected defects, the data quality comprising a pocket of defect concentration; and   transmitting, to a client device, information related to the determined pocket of defect concentration to a client device.   
     
     
         6 . The method of  claim 5 , further comprising:
 clustering the data elements into multiple segments based on the defects.   
     
     
         7 . The method of  claim 5 , wherein obtaining the set of rules includes:
 generating a first set of rules based on a property of the set of data elements; and   generating a second set of rules based on the set of data elements and the first set of rules.   
     
     
         8 . The method of  claim 7 , wherein assessing the data elements includes:
 assessing the data elements based on the second set of rules.   
     
     
         9 . The method of  claim 7 , wherein the first set of rules includes a null rule, a range rule, a uniqueness rule, a valid value rule, a format rule, a conditional rule, or a consistency rule. 
     
     
         10 . The method of  claim 7 , wherein the second set of rules includes a framework criteria comprising a support value, a confidence value, or a lift ratio. 
     
     
         11 . The method of  claim 7 , wherein obtaining the set of rules includes:
 generating the second set of rules based on a user input for adjusting the first set of rules.   
     
     
         12 . The method of  claim 7 , wherein obtaining the set of rules includes:
 generating the first set of rules based on a property of the set of data elements.   
     
     
         13 . The method of  claim 12 , wherein the property includes a portion of a name of the first data elements. 
     
     
         14 . The method of  claim 5 , wherein obtaining the set of rules includes:
 generating a data profile based on the set of data elements; and   generating a first set of rules based on the data profile.   
     
     
         15 . The method of  claim 14 , wherein the data profile includes organizational information, authentication timestamp information, network node location, network node preference, access point location, or access point preference. 
     
     
         16 . The method of  claim 5 , wherein extracting the data elements includes parsing the data elements according to an alphanumeric identifier or a data list from the data source. 
     
     
         17 . A non-transitory computer-readable medium for providing data quality management, comprising instructions that, when executed by one or more processors, cause operations comprising:
 obtaining set of rules for assessing data quality of a set of data elements;   extracting a plurality of data elements from a data source;   assessing the data elements based on a comparison of the data elements to the set of rules;   detecting a plurality of defects based on the comparison;   determining data quality according to the detected defects, the data quality comprising a pocket of defect concentration; and   transmitting, to a client device, information related to the determined pocket of defect concentration to a client device.   
     
     
         18 . The computer-readable medium of  claim 17 , wherein the operations further comprise:
 clustering the data elements into multiple segments based on the defects.   
     
     
         19 . The computer-readable medium of  claim 17 , wherein obtaining the set of rules includes:
 generating a first set of rules based on a property of the set of data elements; and   generating a second set of rules based on the set of data elements and the first set of rules.   
     
     
         20 . The computer-readable medium of  claim 19 , wherein assessing the data elements includes:
 assessing the data elements based on the second set of rules.

Join the waitlist — get patent alerts

Track US2021279215A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.