US2023114836A1PendingUtilityA1

High-speed scanning parser for scalable collection of statistics and use in preparing data for machine learning

Assignee: IQVIA INCPriority: May 10, 2019Filed: Dec 14, 2022Published: Apr 13, 2023
Est. expiryMay 10, 2039(~12.8 yrs left)· nominal 20-yr term from priority
G06F 16/254G06N 20/00G06F 2212/7202G06F 9/544G06F 12/0284
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A parser is deployed early in a machine learning pipeline to read raw data and collect useful statistics about the raw data's content to determine which items of raw data exhibit a proxy for feature importance for the machine learning model. The parser operates at high speeds that approach the disk's absolute throughput while utilizing a small memory footprint. Utilization of the parser enables the machine learning pipeline to receive a fraction of the total raw data that would otherwise be available. Several scans through the data are performed, by which proxies for feature importance are indicated and irrelevant features may be discarded and thereby not forwarded to the machine learning pipeline. This reduces the amount of memory and other hardware resources used at the server and also expedites the machine learning process.

Claims

exact text as granted — not AI-modified
What is claimed: 
     
         1 . A computer-implemented method comprising:
 allocating a buffer corresponding to one or more sizes of data in a set of data;   transferring the set of data into the allocated buffer;   parsing the set of data in the allocated buffer to identify a first set of features and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and   training a machine-learning system using the identified subset of data.   
     
     
         2 . The method of  claim 1 , wherein the catalogue comprises data characteristics to identify the first set of features. 
     
     
         3 . The method of  claim 1 , further comprising:
 determining whether the first set of features exceed a threshold level.   
     
     
         4 . The method of  claim 1 , further comprising:
 utilizing the catalogue to reserve memory for processing of the set of data.   
     
     
         5 . The method of  claim 1 , wherein a utilization of online statistics of the first set of features reduce memory allocations. 
     
     
         6 . The method of  claim 1 , further comprising:
 utilizing the catalogue to identify whether the first set of data features or second set of data features exceed a threshold level.   
     
     
         7 . The method of  claim 1 , wherein the catalogue is produced based on one or more scans of the set of data. 
     
     
         8 . A computer program product comprising a tangible storage medium encoded with processor-readable instructions that, when executed by one or more processors, enable the computer program product to:
 allocate a buffer corresponding to one or more sizes of data in a set of data;   transfer the set of data into the allocated buffer;   parse the set of data in the allocated buffer to identify a first set of features and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and   train a machine-learning system using the identified subset of data.   
     
     
         9 . The computer program product of  claim 8 , wherein the catalogue is used to identify the first set of features to be placed into the machine-learning system. 
     
     
         10 . The computer program product of  claim 8 , wherein the catalogue enables a parser to determine if the first set of features exceed a threshold level. 
     
     
         11 . The computer program product of  claim 8 , wherein the catalogue production includes collecting statistics within the set of data. 
     
     
         12 . The computer program product of  claim 8 , wherein contents of the catalogue are accessed in relation to processing capability. 
     
     
         13 . The computer program product of  claim 8 , wherein the catalogue production creates an allocated memory footprint. 
     
     
         14 . The computer program product of  claim 8 , wherein the set of data is labeled within the catalogue. 
     
     
         15 . A computer system connected to a network, the system comprising:
 one or more processors configured to:
 allocate a buffer corresponding to one or more sizes of data in a set of data; 
 transfer the set of data into the allocated buffer; 
 parse the set of data in the allocated buffer to identify a first set of features 
   
       and a second set of features within the set of data, wherein the parsing of the set of data produces a catalogue, and wherein a subset of the set of data is identified based on the first set features; and
 train a machine-learning system using the identified subset of data. 
 
     
     
         16 . The computer system of  claim 15 , wherein a determination is made as to whether the first set of features exceed a threshold level. 
     
     
         17 . The computer system of  claim 15 , wherein the catalogue is used to identify differences between the first set of features and the second set of features. 
     
     
         18 . The computer system of  claim 15 , wherein the catalogue production reduces one or more memory allocations. 
     
     
         19 . The computer system of  claim 15 , wherein the catalogue production enables statistics and/or one or more flags within the set of data to be collected. 
     
     
         20 . The computer system of  claim 15 , wherein the parser identifies when the first or second set of features pass a threshold level.

Join the waitlist — get patent alerts

Track US2023114836A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.