Systems and methods for preparing data for use by machine learning algorithms
Abstract
Historical data used to train machine learning algorithms can have thousands of records with hundreds of fields, and inevitably includes faulty data that affects the accuracy and utility of a primary model machine learning algorithm. To improve dataset integrity it is segregated into a clean dataset having no invalid data values and a faulty dataset having the invalid data values. The clean dataset is used to produce a secondary model machine learning algorithm trained to generate from plural complete data records a replacement value for a single invalid data value in a data record, and a tertiary model machine learning clustering algorithm trained to generate from plural complete data records replacement values for multiple invalid data values. Substituting the replacement data values for invalid data values in the faulty dataset creates augmented training data which is combined with clean data to train a more accurate and useful primary model.
Claims
exact text as granted — not AI-modified1 . In a system for preparing a plurality of historical data records for use in training a primary model machine learning algorithm, wherein each historical data record includes a plurality of fields containing data values designated as inputs for training the primary model machine learning algorithm to generate an output of interest, a computer-implemented method of preparing the designated inputs in the plurality of historical data records in order to increase the utility and accuracy of the trained primary model machine learning algorithm when the designated inputs in the historical data records include invalid data values, the method comprising:
segregating a base dataset containing the historical data records into a faulty dataset having incomplete data records with invalid data values and a clean dataset having complete data records with no invalid data values; storing the faulty dataset and the clean dataset in a computer memory; producing from the stored clean dataset at least one of (i) plural computer-implemented secondary model machine learning algorithms trained to generate from values of the respective fields designated as inputs in plural complete data records a replacement value for a single invalid data value in a corresponding field designated as an input in an incomplete data record, and (ii) a computer-implemented tertiary model machine learning clustering algorithm trained to generate from plural complete data records comprising all values of fields designated as inputs replacement values for multiple invalid data values designated as inputs in an incomplete data record; and using a computer-implemented program to create augmented training data records by substituting the replacement data values for at least some of the respective invalid data values designated as inputs in the stored faulty dataset, whereby said augmented training data records can be used with complete data records in the clean dataset to train the primary model machine learning algorithm to improve the accuracy thereof when generating an output of interest from a new data record.
2 . A system as in claim 1 , wherein the method further comprises:
training the primary model machine learning algorithm using the augmented training data records; obtaining a new data record with fields corresponding to respective ones of the fields in the historical data records designated as inputs; storing the new data record in a computer memory; completing the new data record by applying to the stored new data record the secondary model machine learning algorithm to generate a replacement value for data in the new data record with a single field containing an invalid data value, and the tertiary model machine learning clustering algorithm to generate replacement values for data in the new data record with multiple fields containing invalid data values; and using the trained primary model machine learning algorithm to generate the output of interest from the new data record.
3 . A system as in claim 1 , wherein at least one field in an historical data record is designated as an output of interest comprising a target data value in numeric form, and the primary machine learning algorithm uses supervised learning to fit a curve relating the data values of fields designated as inputs in the historical data records to the numeric target data values in the historical data records.
4 . A system as in claim 1 , wherein at least one field in an historical data record is designated as an output of interest comprising target data values in the form of two or more discrete classes, and the primary machine learning algorithm uses supervised learning to maximize the probability that the values of the data designated as inputs in the historical data records determine that the data record is a member of one of the two or more discrete classes comprising the target data values in the historical data records.
5 . A system as in claim 1 , wherein the primary machine learning algorithm output of interest is an identification of a collection of data records whose values designated as inputs are more similar to other data records in the collection than said input values are similar to values designated as inputs in data records which are not in the collection.
6 . A system as in claim 1 , wherein the secondary model machine learning algorithm comprises at least one of a prediction model for generating replacement values for fields having values in continuous numeric form and a classification model for generating replacement values for fields having values in the form of discrete classes.
7 . A system as in claim 1 , wherein the secondary model machine learning algorithm comprises a multi-layer feed-forward neural network trained by back-propagation.
8 . A system as in claim 1 , wherein the tertiary model machine learning algorithm comprises a self-organizing map characterized by a plurality of clusters based on the total number of historical data records.
9 . A system as in claim 1 , wherein the method of preparing the designated inputs in the plurality of historical data records for training of the primary model further comprises:
using a heuristic analysis to identify any of the fields designated as inputs for the primary model machine learning algorithm that contain data values having no utility for training the primary machine learning algorithm to generate the output of interest; creating a reduced clean dataset without the fields containing data values identified as having no utility for training the primary machine learning algorithm to generate the output of interest and storing the reduced clean dataset in a computer memory; creating an auxiliary clean dataset without any fields representing the output of interest and storing the auxiliary clean data set; and using the stored auxiliary clean dataset as training data for the plural secondary model machine learning algorithms and the tertiary model machine learning clustering algorithm.
10 . A method of using a computer-implemented primary model machine learning algorithm trained with a plurality of historical data records, wherein each historical data record includes a plurality of fields designated as inputs for training the primary model machine learning algorithm to generate an output of interest, wherein the method generates a corresponding output of interest from a new data record with a plurality of fields corresponding to the fields in the historical data records designated as inputs when one or more of the fields in the new data record contains an invalid data value, the method comprising:
using one of a computer-implemented secondary model machine learning algorithm trained using the historical data records to generate a replacement value for a single field containing an invalid data value, and a computer-implemented tertiary model machine learning clustering algorithm trained using the historical data records to generate replacement values for a data record with multiple fields containing invalid data values; completing the new data record by substituting the one or more replacement values corresponding to the data values in respective fields of the new data record containing an invalid data value; and using the primary model machine learning algorithm to generate from the completed new data record the output of interest associated with the new data record.
11 . A method as in claim 10 , further comprising:
accessing a base dataset with a plurality of the historical data records; segregating from the stored base dataset a clean dataset having complete historical data records with no invalid data values; storing the clean dataset in a computer memory; producing from the stored clean dataset the secondary model machine learning algorithm and the tertiary model machine learning clustering algorithm.
12 . A method as in claim 11 , further comprising;
storing in a computer memory a faulty dataset having incomplete historical data records with invalid data values; using a computer-implemented program to create augmented training data records by substituting the replacement data values for at least some of the respective invalid data values in data records in the stored faulty dataset and storing the augmented training data records in a computer memory; and training the primary model machine learning algorithm using the data records in the clean dataset combined with the augmented training data records.
13 .- 26 . (canceled)Join the waitlist — get patent alerts
Track US2020401939A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.