High-speed scanning parser for scalable collection of statistics and use in preparing data for machine learning
Abstract
A parser is deployed early in a machine learning pipeline to read raw data and collect useful statistics about the raw data's content to determine which items of raw data exhibit a proxy for feature importance for the machine learning model. The parser operates at high speeds that approach the disk's absolute throughput while utilizing a small memory footprint. Utilization of the parser enables the machine learning pipeline to receive a fraction of the total raw data that would otherwise be available. Several scans through the data are performed, by which proxies for feature importance are indicated and irrelevant features may be discarded and thereby not forwarded to the machine learning pipeline. This reduces the amount of memory and other hardware resources used at the server and also expedites the machine learning process.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . A computer-implemented method comprising:
allocating a buffer corresponding to one or more sizes of data in a set of data; transferring the set of data into the allocated buffer; parsing the set of data in the allocated buffer to identify a first set of features and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and training a machine-learning system using the identified subset of data.
2 . The method of claim 1 , wherein the catalogue comprises data characteristics to identify the first set of features.
3 . The method of claim 1 , further comprising:
determining whether the first set of features exceed a threshold level.
4 . The method of claim 1 , further comprising:
utilizing the catalogue to reserve memory for processing of the set of data.
5 . The method of claim 1 , wherein a utilization of online statistics of the first set of features reduce memory allocations.
6 . The method of claim 1 , further comprising:
utilizing the catalogue to identify whether the first set of data features or second set of data features exceed a threshold level.
7 . The method of claim 1 , wherein the catalogue is produced based on one or more scans of the set of data.
8 . A computer program product comprising a tangible storage medium encoded with processor-readable instructions that, when executed by one or more processors, enable the computer program product to:
allocate a buffer corresponding to one or more sizes of data in a set of data; transfer the set of data into the allocated buffer; parse the set of data in the allocated buffer to identify a first set of features and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and train a machine-learning system using the identified subset of data.
9 . The computer program product of claim 8 , wherein the catalogue is used to identify the first set of features to be placed into the machine-learning system.
10 . The computer program product of claim 8 , wherein the catalogue enables a parser to determine if the first set of features exceed a threshold level.
11 . The computer program product of claim 8 , wherein the catalogue production includes collecting statistics within the set of data.
12 . The computer program product of claim 8 , wherein contents of the catalogue are accessed in relation to processing capability.
13 . The computer program product of claim 8 , wherein the catalogue production creates an allocated memory footprint.
14 . The computer program product of claim 8 , wherein the set of data is labeled within the catalogue.
15 . A computer system connected to a network, the system comprising:
one or more processors configured to:
allocate a buffer corresponding to one or more sizes of data in a set of data;
transfer the set of data into the allocated buffer;
parse the set of data in the allocated buffer to identify a first set of features
and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and
train a machine-learning system using the identified subset of data.
16 . The computer system of claim 15 , wherein a determination is made as to whether the first set of features exceed a threshold level.
17 . The computer system of claim 15 , wherein the catalogue is used to identify differences between the first set of features and the second set of features.
18 . The computer system of claim 15 , wherein the catalogue production reduces one or more memory allocations.
19 . The computer system of claim 15 , wherein the catalogue production enables statistics and/or one or more flags within the set of data to be collected.
20 . The computer system of claim 15 , wherein the parser identifies when the first or second set of features pass a threshold level.Join the waitlist — get patent alerts
Track US2023114836A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.