Systems and methods for data validation and transformation of data in a zero-trust environment
Abstract
Systems and methods for the validation and transform of data for processing by an algorithm is provided. In some embodiments, input data is cleaned, and then the domain of the data is determined. The domain of the data refers to the data type. A validation of the data occurs. The validation is for the ranges and distribution that the data should have, according to the domain, versus the actual data ranges and distribution. Data that fails the validation undergo a transform step and then are re-validated. This process is iterative until the data set passes validation. Transforms are first selected based upon the data domain that was determined prior. Transforms that fit a range requirement, or a distribution type may be selected. In alternate embodiments, machine learning (ML) may be employed to train models, exclusive to a given domain, to identify needed transforms.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computerized method of processing input data comprising:
identifying a domain for a set of input data; validating the set of input data by range and distribution responsive to the domain; when the validating fails, transforming the input data by at least one of a range transformation, a distribution transformation and a machine learning (ML) transformation; iteratively validating and transforming the input data until validation passes; and processing the validated data using at least one algorithm.
2 . The method of claim 1 , wherein the domain is at least one of pathology dependent and financial use case dependent.
3 . The method of claim 1 , further comprising cleaning the input data.
4 . The method of claim 1 , wherein the ML transform is trained on domain specific datasets.
5 . The method of claim 1 , wherein the validation includes comparing the datasets to an expected range and distribution curve for the data domain.
6 . The method of claim 5 , wherein the validation of expected distribution is a curve which fits within two standard deviations of the expected distribution.
7 . The method of claim 5 , wherein the validation of expected distribution is a curve which fits within a configurable threshold of standard deviations of the expected distribution.
8 . A computerized system of processing input data comprising:
a computer server for identifying a domain for a set of input data using an AI clustering model, validating the set of input data by range and distribution responsive to the domain using a statistical model, and when the validating fails, transforming the input data by at least one of a range transformation, a distribution transformation and a machine learning (ML) transformation, and iteratively validating and transforming the input data until validation passes, and processing the validated data using at least one algorithm.
9 . The system of claim 8 , wherein the domain is pathology dependent.
10 . The system of claim 8 , further comprising a database for cleaning and storing the input data.
11 . The system of claim 8 , wherein the ML transform is trained on domain specific datasets.
12 . The system of claim 8 , wherein the validation includes comparing the datasets to an expected range and distribution curve for the data domain.
13 . The system of claim 12 , wherein the validation of expected distribution is a curve which fits within two standard distributions of the expected distribution.
14 . The system of claim 12 , wherein the validation of expected distribution is a curve which fits within one standard distributions of the expected distribution.Join the waitlist — get patent alerts
Track US2023205917A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.