US2013218908A1PendingUtilityA1

Computing and applying order statistics for data preparation

Individually held — no corporate assignee on recordPriority: Feb 17, 2012Filed: Feb 17, 2012Published: Aug 22, 2013
Est. expiryFeb 17, 2032(~5.6 yrs left)· nominal 20-yr term from priority
G06Q 30/06G06Q 10/10G06F 16/27G06F 16/11
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Provided are techniques for generating order statistics and error bounds. For each of multiple, distributed data sources, a finite number of data bins are created for each field in that data source. Data values in each of the multiple, distributed data sources are processed to generate basic summaries for each of the data bins in a single pass of the data values. The data bins from each of the multiple, distributed data sources are sorted. One or more approximate order statistics are computed for a data set by accumulating counts from a number of the sorted data bins. Lower and upper error bounds are provided for each of the computed one or more approximate order statistics, wherein the lower and upper error bounds are values delimiting an interval containing a true value of an order statistic.

Claims

exact text as granted — not AI-modified
1 . A computer system for generating order statistics and error bounds, comprising:
 a processor; and   a storage device connected to the processor, wherein the storage device has stored thereon a program, and wherein the processor is configured to execute instructions of the program to perform operations, wherein the operations comprise:
 for each of multiple, distributed data sources, creating a finite number of data bins for each field in that data source; 
 processing data values in each of the multiple, distributed data sources to generate basic summaries for each of the data bins in a single pass of the data values; 
 sorting the data bins from each of the multiple, distributed data sources; 
 computing one or more approximate order statistics for a data set by accumulating counts from a number of ordered data bins; 
 providing lower and upper error bounds for each of the computed one or more approximate order statistics, wherein the lower and upper error bounds are values delimiting the interval containing the true value of an order statistic. 
   
     
     
         2 . The computer system of  claim 1 , wherein the basic summaries for a data bin comprises a count, a mean, a lower bound, and an upper bound for that data bin. 
     
     
         3 . The computer system of  claim 1 , further comprising calculating a power transformation parameter for a Box-Cox transformation using the computed one or more approximate order statistics. 
     
     
         4 . The computer system of  claim 1 , further comprising:
 for each of the finite number of data bins, generating a data bin of zero width; and   in response to receiving a new data value,
 determining whether the new data value is to be added to an existing data bin; 
 in response to determining that the new data value is to be added to the existing data bin, 
 adding the new data value to the existing data bin; and 
 updating basic summaries of the existing data bin; 
 in response to determining that the new data value is not to be added to the existing bin, creating a new data bin for the new data value; and 
 creating basic summaries for the new data bin. 
   
     
     
         5 . The computer system of  claim 4 , further comprising:
 merging each new data bin with the existing bins in batches by adjusting the basic summaries of each data bin involved in a merge when the number of bins exceeds a preset threshold and while ensuring that width of the merged bins does not exceed an approximation bound.   
     
     
         6 . The computer system of  claim 1 , wherein a width of each data bin is maintained within limits bounded by a range of data values divided by the finite number of data bins. 
     
     
         7 . A computer program product for generating order statistics and error bounds, the computer program product comprising:
 a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code comprising:   computer readable program code, when executed by a processor of a computer, configured to perform:
 for each of multiple, distributed data sources, creating a finite number of data bins for each field in that data source; 
 processing data values in each of the multiple, distributed data sources to generate basic summaries for each of the data bins in a single pass of the data values; 
 sorting the data bins from each of the multiple, distributed data sources; 
 computing one or more approximate order statistics for a data set by accumulating counts from a number of the sorted data bins; 
 providing lower and upper error bounds for each of the computed one or more approximate order statistics, wherein the lower and upper error bounds are values delimiting the interval containing the true value of an order statistic. 
   
     
     
         8 . The computer program product of  claim 7 , wherein the basic summaries for a data bin comprises a count, a mean, a lower bound, and an upper bound for that data bin. 
     
     
         9 . The computer program product of  claim 7 , further comprising calculating a power transformation parameter for a Box-Cox transformation using the computed one or more approximate order statistics. 
     
     
         10 . The computer program product of  claim 7 , further comprising:
 for each of the finite number of data bins, generating a data bin of zero width; and   in response to receiving a new data value,
 determining whether the new data value is to be added to an existing data bin; 
 in response to determining that the new data value is to be added to the existing data bin, 
 adding the new data value to the existing data bin; and 
 updating basic summaries of the existing data bin; 
 in response to determining that the new data value is not to be added to the existing bin, creating a new data bin for the new data value; and 
 creating basic summaries for the new data bin. 
   
     
     
         11 . The computer program product of  claim 10 , further comprising:
 merging each new data bin with the existing bins in batches by adjusting the basic summaries of each data bin involved in a merge when the number of bins exceeds a preset threshold and while ensuring that width of the merged bins does not exceed an approximation bound.   
     
     
         12 . The computer program product of  claim 7 , wherein a width of each data bin is maintained within limits bounded by a range of data values divided by the finite number of data bins.

Join the waitlist — get patent alerts

Track US2013218908A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.