US2022229809A1PendingUtilityA1

Method and system for flexible, high performance structured data processing

Assignee: ANDITI PTY LTDPriority: Jun 28, 2016Filed: Apr 4, 2022Published: Jul 21, 2022
Est. expiryJun 28, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G06F 9/5027G06F 16/29G06F 16/258G06F 16/13
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Described herein is a method and system for flexible, high performance structured data processing. The method and system contains techniques for balancing and jointly optimising processing speed, resource utilisation, flexibility, scalability, and configurability in one workflow. A prime example of its application is the analysis of spatial data, e.g. LiDAR and imagery. However, the invention is applicable to a wide range of structured data problems in a variety of dimensions and settings.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A method of allocating computer resources to a data processing operation for processing structured data stored as a dataset, the method comprising:
 pre-processing the dataset to generate a metadata file including characteristics of the dataset;   dividing the dataset into a plurality of work units, each work unit indicative of a subset of the data contained in the dataset;   creating a list of work units of a predetermined size based on the characteristics of the dataset;   calculating the computational complexity of each work unit based on the size of work units and characteristics of the dataset;   determining memory requirements for each work unit;   determining available memory of connected computer nodes; and   allocating work units to connected computer nodes for processing based on available memory and the number of processes running.   
     
     
         2 . The method of  claim 1 , further comprising the step of merging or subdividing work units based on the available memory within one or more of the connected computer nodes. 
     
     
         3 . The method of  claim 1 , wherein the pre-processing further comprises generating a reference file indicating one or more predetermined characteristics of the data that are contained within each of the discrete data files. 
     
     
         4 . The method of  claim 1 , wherein the step of pre-processing the dataset occurs in conjunction with indexing the dataset. 
     
     
         5 . The method of  claim 3 , wherein the predetermined characteristics include data bounds and associated file names for each of the discrete data files in the dataset. 
     
     
         6 . The method of  claim 1 , wherein pre-processing further comprises pre-classifying the discrete data files to calculate one or more data metrics. 
     
     
         7 . The method of  claim 6 , wherein the data metrics include a determination of the likelihood of the presence of certain data features in discrete data files. 
     
     
         8 . The method of  claim 1  wherein the step of pre-processing the dataset includes the steps of:
 i) opening each discrete data file; 
 ii) determining the data bounds for each discrete data file; and 
 iii) storing the determined data bounds and an associated filename for each discrete data file in the reference file. 
 
     
     
         9 . The method of  claim 8 , wherein the dataset includes spatial data. 
     
     
         10 . The method of  claim 11  wherein the spatial data includes imagery data. 
     
     
         11 . The method of  claim 10  wherein the structured dataset is a dataset of time series data. 
     
     
         12 . A method according to  claim 6 , wherein the pre-classifying comprises:
 a1) creating a metadata file;   a2) opening the discrete data files;   a3) dividing the dataset into predefined data cells and determining at least one data metric for each of the data cells; and   a4) storing the at least one data metric in the metadata file in association with an associated data cell identifier for each data cell and an identifier of the discrete data file(s) associated with each data cell.   
     
     
         13 . The method of  claim 12  wherein the at least one data metric comprises a measure of likelihood that the data of an individual data file includes specific spatial, temporal or spectral features. 
     
     
         14 . The method of  claim 13  wherein the at least one data metric comprises a measure of quality of the data within an individual data file. 
     
     
         15 . The method of  claim 1 , further comprising the steps of defining plugin connections for processing the dataset; dynamically allocating data processing tasks to connected computer nodes; performing data processing on the selection of data; and generating an output of the data processing of the selection of data.

Join the waitlist — get patent alerts

Track US2022229809A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.