US2022171781A1PendingUtilityA1

System And Method For Analyzing Data Records

Assignee: GOOGLE LLCPriority: Jun 18, 2004Filed: Feb 16, 2022Published: Jun 2, 2022
Est. expiryJun 18, 2024(expired)· nominal 20-yr term from priority
G06F 11/1482G06F 16/24561G06F 16/2471G06F 16/285Y10S707/99933Y10S707/99937G06F 16/1858
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for analyzing input data records are provided in which a master process initiates a plurality of concurrent first processes each of which comprises, for each data record in at least a subset of a plurality of input data records, creating a parsed representation of the data record and independently applying a procedural language query to the parsed representation to extract one or more values. A respective emit operator is applied to at least one of the extracted one or more values thereby adding corresponding information to a respective intermediate data structure. The respective emit operator implements one of a predefined set of statistical information processing functions. The master process also initiates a plurality of second processes each of which aggregates information from a corresponding subset of intermediate data structures to produce aggregated data that is, in turn, combined to produce output data.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A computer-implemented method of analyzing a plurality of data records, comprising:
 executing a master process that forms a procedure comprising:   (A) initiating a plurality of first processes, wherein each respective process of the plurality of first processes is executed concurrently by the master process and comprises:   for each data record in at least a subset of the plurality of data records:
 splitting the plurality of data records into a plurality of data blocks; 
 independently applying map functions to each of the plurality of data blocks to generate intermediate data; and 
 partitioning the intermediate data between a plurality of intermediate files distributed across a plurality of local databases; 
   (B) initiating a plurality of second processes, wherein each respective process of the plurality of second processes aggregates information from a subset of the intermediate data to produce aggregated data;   wherein the computer implemented method further combines the produced aggregated data to produce output data.   
     
     
         2 . The method of  claim 1 , wherein each process in the plurality of first processes is compiled into a binary file prior to execution. 
     
     
         3 . The method of  claim 1 , wherein each process in the plurality of first processes is implemented as a user defined object in accordance with an object oriented programming technique. 
     
     
         4 . The method of  claim 1 , wherein each process in the plurality of first processes is implemented as a user defined object that is derived from a base class in accordance with an object oriented programming technique. 
     
     
         5 . The method of  claim 1 , wherein the initiating the plurality of first process comprises creating a respective object in a plurality of object for each first process in the plurality of first processes. 
     
     
         6 . The method of  claim 1 , wherein the independently applying map functions are executed at the plurality of local databases. 
     
     
         7 . The method of  claim 6 , wherein a number of the map functions applied to the plurality of data blocks is sufficient to store all intermediate data generated by the map functions at the plurality of local databases. 
     
     
         8 . The method of  claim 1 , wherein a second process in the plurality of second processes comprises one or more of the following: a function for counting occurrences of distinct values in the corresponding subset of intermediate data structures, a maximum value function for identifying a maximum value in the corresponding subset of intermediate data structures, a minimum value function for identifying a minimum value in the corresponding subset of intermediate data structures, a statistical sampling function for applying a statistical function to the corresponding subset of intermediate data structures, a function for identifying values that occur most frequently in the corresponding subset of intermediate data structures, and a function for estimating a total number of unique values in the corresponding subset of intermediate data structures. 
     
     
         9 . The method of  claim 1 , wherein the intermediate data is structured as key-value pairs. 
     
     
         10 . The method of  claim 1 , wherein the intermediate data comprises a table having at least one index whose index values comprise one or more values extracted from the plurality of data blocks by the map functions. 
     
     
         11 . The method of  claim 10 , wherein the aggregation of information from the subset of the intermediate data structures to produce the aggregated data combines the extracted one or more values having the same index values. 
     
     
         12 . The method of  claim 1 , further comprising executing the plurality of second processes in parallel. 
     
     
         13 . The method of  claim 1 , wherein the plurality of data records comprises one or more of the following types of data records: log files, transaction records, and documents. 
     
     
         14 . The method of  claim 1 , wherein the intermediate data comprises a table having a plurality of indices, wherein each of the plurality of indices is dynamically generated in accordance with a corresponding one or more values extracted from the plurality of data blocks by the map functions. 
     
     
         15 . A computer system with one or more processors and memory for analyzing a plurality of data records, the computer system comprising memory and one or more processors, wherein the memory stores instructions for executing a master process that forms a procedure comprising:
 (A) initiating a plurality of first processes, wherein each respective process of the plurality of first processes is executed concurrently by the master process and comprises:   for each data record in at least a subset of the plurality of data records:
 splitting the plurality of data records into a plurality of data blocks; 
 independently applying map functions to each of the plurality of data blocks to generate intermediate data; and 
 partitioning the intermediate data between a plurality of intermediate files distributed across a plurality of local databases; 
   (B) initiating a plurality of second processes, wherein each respective process of the plurality of second processes aggregates information from a subset of the intermediate data to produce aggregated data;   wherein the computer implemented method further combines the produced aggregated data to produce output data.   
     
     
         16 . The computer system of  claim 15 , wherein each first process in the plurality of first processes is implemented as a user defined object in accordance with an object oriented programming technique. 
     
     
         17 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs for executing a master process that forms a procedure, which when executed by a computer system, causes the computer system to:
 (A) initiate a plurality of first processes, wherein each respective process of the plurality of first processes is executed concurrently by the master process and comprises:   for each data record in at least a subset of the plurality of data records:
 split the plurality of data records into a plurality of data blocks; 
 independently apply map functions to each of the plurality of data blocks to generate intermediate data; and 
 partition the intermediate data between a plurality of intermediate files distributed across a plurality of local databases; 
   (B) initiate a plurality of second processes, wherein each respective process of the plurality of second processes aggregates information from a subset of the intermediate data to produce aggregated data;   wherein the computer implemented method further combines the produced aggregated data to produce output data.   
     
     
         18 . The non-transitory computer readable storage medium of  claim 17 , wherein each first process in the plurality of first processes is implemented as a user defined object in accordance with an object oriented programming technique. 
     
     
         19 . The computer system of  claim 17 , wherein the initiating the plurality of first processes comprises creating a respective object in a plurality of object for each first process in the plurality of first processes. 
     
     
         20 . The non-transitory computer readable storage medium of  claim 19 , wherein the initiating the plurality of first processes comprises creating a respective object in a plurality of objects for each first process in the plurality of first processes.

Join the waitlist — get patent alerts

Track US2022171781A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.