Data ingestion to generate layered dataset interrelations to form a system of networked collaborative datasets
Abstract
Various embodiments relate generally to data science and data analysis, and computer software and systems to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby data ingestion is performed to form data representing layered data files and data arrangements to facilitate, for example, interrelations among a system of networked collaborative datasets. In some examples, a method may include forming a first layer data file and a second layer data file, assigning addressable identifiers to uniquely identify units of data and data units to facilitate the linking of data, and implementing selectively one or more of a unit of data and a data unit as a function of a context of a data access request for a collaborative dataset.
Claims
exact text as granted — not AI-modified1 . A method, comprising:
receiving data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format; adapting the initial data format to the set of data to form a dataset having a first data format; forming a multiple layer data file to link other data in a second data format to the dataset in the first data format; interpreting a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset; associating the data subset with annotative data identifying the inferred attribute; and converting the dataset in the first data format to an atomized dataset using another data format.
2 . The method of claim 1 , further comprising associating the dataset in the first data format with an identifier.
3 . The method of claim 1 , wherein the data subset comprises a columnar representation of data in a tabular format.
4 . The method of claim 1 , wherein the data subset comprises a columnar representation of data, the columnar representation of data being annotated.
5 . The method of claim 1 , further comprising forming a collaborative dataset by retrieving another atomized dataset.
6 . The method of claim 1 , further comprising:
correlating the data subset over another dataset; and inferring the data classification of the data subset by analyzing one or more data strings.
7 . The method of claim 1 , wherein the atomized dataset comprises a set of atomized datapoints.
8 . The method of claim 1 , wherein the atomized dataset comprises a set of atomized datapoints, the set of atomized datapoints comprising a triple.
9 . The method of claim 1 , wherein forming the multiple layer data files comprises
forming a first layer data file to reference the other data in the second data format in which a unit of data in the other data in the second data format is configured to link with another data file; and forming a second layer data file that includes a subset of data based on the other data in the second data format, in which a unit of data in the subset of data in the second data format being configured to link to a unit of data in the first data format.
10 . The method of claim 1 , further comprising
assigning an addressable identifier to uniquely identify a unit of data to facilitate linking data between the dataset in the first format and the other data in the second data format, at least one of the addressable identifiers references a triplestore database; and implementing selectively the unit of data as a function of context of a data access request.
11 . A system, comprising:
a memory including executable instructions; and a processor responsive to executing instructions, is configured to: receive data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format; adapt the initial data format to the set of data to form a dataset having a first data format; form a multiple layer data file to link other data in a second data format to the dataset in the first data format; interpret a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset; associate the data subset with annotative data identifying the inferred attribute; and convert the dataset in the first data format to an atomized dataset using another data format.
12 . The system of claim 11 , wherein the processor is further configured to associate the dataset in the first data format with an identifier.
13 . The system of claim 11 , wherein the data subset further comprises a columnar representation of data in a tabular format.
14 . The system of claim 11 , wherein the data subset comprises a columnar representation of data, the columnar representation of the data being annotated.
15 . The system of claim 11 , wherein the processor is further configured to form a collaborative dataset by retrieving another atomized dataset.
16 . The system of claim 11 , wherein the processor is further configured to:
correlate the data subset over another dataset; and infer the data classification of the data subset by analyzing one or more data strings.
17 . The system of claim 11 , wherein the atomized dataset comprises a set of atomized datapoints, the set of atomized datapoints comprising a triple.
18 . The system of claim 11 , wherein the processor is further configured to:
form a first layer data file to reference the other data in the second data format in which a unit of data in the other data in the second data format is configured to link with another data file; and form a second layer data file that includes a subset of data based on the other data in the second data format, in which a unit of data in the subset of data in the second data format being configured to link to a unit of data in the first data format.
19 . The system of claim 1 , wherein the processor is further configured to:
assign an addressable identifier to uniquely identify the units of data and the data units to facilitate linking data between the dataset in the first format and the other data in the second data format, at least one of the addressable identifiers references a triplestore database; and implement selectively one or more of a unit of data and a data unit as a function of context of a data access request.
20 . A non-transitory computer readable medium having one or more computer program instructions configured to perform a method, the method comprising:
receiving data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format; adapting the initial data format to the set of data to form a dataset having a first data format; forming a multiple layer data file to link other data in a second data format to the dataset in the first data format; interpreting a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset; associating the data subset with annotative data identifying the inferred attribute; and converting the dataset in the first data format to an atomized dataset using another data format.Join the waitlist — get patent alerts
Track US2023109821A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.