US2022229809A1PendingUtilityA1
Method and system for flexible, high performance structured data processing
Est. expiryJun 28, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G06F 9/5027G06F 16/29G06F 16/258G06F 16/13
50
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Described herein is a method and system for flexible, high performance structured data processing. The method and system contains techniques for balancing and jointly optimising processing speed, resource utilisation, flexibility, scalability, and configurability in one workflow. A prime example of its application is the analysis of spatial data, e.g. LiDAR and imagery. However, the invention is applicable to a wide range of structured data problems in a variety of dimensions and settings.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method of allocating computer resources to a data processing operation for processing structured data stored as a dataset, the method comprising:
pre-processing the dataset to generate a metadata file including characteristics of the dataset; dividing the dataset into a plurality of work units, each work unit indicative of a subset of the data contained in the dataset; creating a list of work units of a predetermined size based on the characteristics of the dataset; calculating the computational complexity of each work unit based on the size of work units and characteristics of the dataset; determining memory requirements for each work unit; determining available memory of connected computer nodes; and allocating work units to connected computer nodes for processing based on available memory and the number of processes running.
2 . The method of claim 1 , further comprising the step of merging or subdividing work units based on the available memory within one or more of the connected computer nodes.
3 . The method of claim 1 , wherein the pre-processing further comprises generating a reference file indicating one or more predetermined characteristics of the data that are contained within each of the discrete data files.
4 . The method of claim 1 , wherein the step of pre-processing the dataset occurs in conjunction with indexing the dataset.
5 . The method of claim 3 , wherein the predetermined characteristics include data bounds and associated file names for each of the discrete data files in the dataset.
6 . The method of claim 1 , wherein pre-processing further comprises pre-classifying the discrete data files to calculate one or more data metrics.
7 . The method of claim 6 , wherein the data metrics include a determination of the likelihood of the presence of certain data features in discrete data files.
8 . The method of claim 1 wherein the step of pre-processing the dataset includes the steps of:
i) opening each discrete data file;
ii) determining the data bounds for each discrete data file; and
iii) storing the determined data bounds and an associated filename for each discrete data file in the reference file.
9 . The method of claim 8 , wherein the dataset includes spatial data.
10 . The method of claim 11 wherein the spatial data includes imagery data.
11 . The method of claim 10 wherein the structured dataset is a dataset of time series data.
12 . A method according to claim 6 , wherein the pre-classifying comprises:
a1) creating a metadata file; a2) opening the discrete data files; a3) dividing the dataset into predefined data cells and determining at least one data metric for each of the data cells; and a4) storing the at least one data metric in the metadata file in association with an associated data cell identifier for each data cell and an identifier of the discrete data file(s) associated with each data cell.
13 . The method of claim 12 wherein the at least one data metric comprises a measure of likelihood that the data of an individual data file includes specific spatial, temporal or spectral features.
14 . The method of claim 13 wherein the at least one data metric comprises a measure of quality of the data within an individual data file.
15 . The method of claim 1 , further comprising the steps of defining plugin connections for processing the dataset; dynamically allocating data processing tasks to connected computer nodes; performing data processing on the selection of data; and generating an output of the data processing of the selection of data.Join the waitlist — get patent alerts
Track US2022229809A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.