US2023109821A1PendingUtilityA1

Data ingestion to generate layered dataset interrelations to form a system of networked collaborative datasets

Assignee: DATA WORLD INCPriority: Jun 19, 2016Filed: Sep 2, 2022Published: Apr 13, 2023
Est. expiryJun 19, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G06F 16/215G06F 16/2471G06F 16/258G06F 16/2423G06F 16/9024
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Various embodiments relate generally to data science and data analysis, and computer software and systems to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby data ingestion is performed to form data representing layered data files and data arrangements to facilitate, for example, interrelations among a system of networked collaborative datasets. In some examples, a method may include forming a first layer data file and a second layer data file, assigning addressable identifiers to uniquely identify units of data and data units to facilitate the linking of data, and implementing selectively one or more of a unit of data and a data unit as a function of a context of a data access request for a collaborative dataset.

Claims

exact text as granted — not AI-modified
1 . A method, comprising:
 receiving data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format;   adapting the initial data format to the set of data to form a dataset having a first data format;   forming a multiple layer data file to link other data in a second data format to the dataset in the first data format;   interpreting a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset;   associating the data subset with annotative data identifying the inferred attribute; and   converting the dataset in the first data format to an atomized dataset using another data format.   
     
     
         2 . The method of  claim 1 , further comprising associating the dataset in the first data format with an identifier. 
     
     
         3 . The method of  claim 1 , wherein the data subset comprises a columnar representation of data in a tabular format. 
     
     
         4 . The method of  claim 1 , wherein the data subset comprises a columnar representation of data, the columnar representation of data being annotated. 
     
     
         5 . The method of  claim 1 , further comprising forming a collaborative dataset by retrieving another atomized dataset. 
     
     
         6 . The method of  claim 1 , further comprising:
 correlating the data subset over another dataset; and   inferring the data classification of the data subset by analyzing one or more data strings.   
     
     
         7 . The method of  claim 1 , wherein the atomized dataset comprises a set of atomized datapoints. 
     
     
         8 . The method of  claim 1 , wherein the atomized dataset comprises a set of atomized datapoints, the set of atomized datapoints comprising a triple. 
     
     
         9 . The method of  claim 1 , wherein forming the multiple layer data files comprises
 forming a first layer data file to reference the other data in the second data format in which a unit of data in the other data in the second data format is configured to link with another data file; and   forming a second layer data file that includes a subset of data based on the other data in the second data format, in which a unit of data in the subset of data in the second data format being configured to link to a unit of data in the first data format.   
     
     
         10 . The method of  claim 1 , further comprising
 assigning an addressable identifier to uniquely identify a unit of data to facilitate linking data between the dataset in the first format and the other data in the second data format, at least one of the addressable identifiers references a triplestore database; and   implementing selectively the unit of data as a function of context of a data access request.   
     
     
         11 . A system, comprising:
 a memory including executable instructions; and   a processor responsive to executing instructions, is configured to:   receive data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format;   adapt the initial data format to the set of data to form a dataset having a first data format;   form a multiple layer data file to link other data in a second data format to the dataset in the first data format;   interpret a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset;   associate the data subset with annotative data identifying the inferred attribute; and   convert the dataset in the first data format to an atomized dataset using another data format.   
     
     
         12 . The system of  claim 11 , wherein the processor is further configured to associate the dataset in the first data format with an identifier. 
     
     
         13 . The system of  claim 11 , wherein the data subset further comprises a columnar representation of data in a tabular format. 
     
     
         14 . The system of  claim 11 , wherein the data subset comprises a columnar representation of data, the columnar representation of the data being annotated. 
     
     
         15 . The system of  claim 11 , wherein the processor is further configured to form a collaborative dataset by retrieving another atomized dataset. 
     
     
         16 . The system of  claim 11 , wherein the processor is further configured to:
 correlate the data subset over another dataset; and   infer the data classification of the data subset by analyzing one or more data strings.   
     
     
         17 . The system of  claim 11 , wherein the atomized dataset comprises a set of atomized datapoints, the set of atomized datapoints comprising a triple. 
     
     
         18 . The system of  claim 11 , wherein the processor is further configured to:
 form a first layer data file to reference the other data in the second data format in which a unit of data in the other data in the second data format is configured to link with another data file; and   form a second layer data file that includes a subset of data based on the other data in the second data format, in which a unit of data in the subset of data in the second data format being configured to link to a unit of data in the first data format.   
     
     
         19 . The system of  claim 1 , wherein the processor is further configured to:
 assign an addressable identifier to uniquely identify the units of data and the data units to facilitate linking data between the dataset in the first format and the other data in the second data format, at least one of the addressable identifiers references a triplestore database; and   implement selectively one or more of a unit of data and a data unit as a function of context of a data access request.   
     
     
         20 . A non-transitory computer readable medium having one or more computer program instructions configured to perform a method, the method comprising:
 receiving data into a collaborative dataset consolidation system, the data being configured to represent a set of data having an initial data format;   adapting the initial data format to the set of data to form a dataset having a first data format;   forming a multiple layer data file to link other data in a second data format to the dataset in the first data format;   interpreting a data subset of the dataset in the first data format against one or more data classifications to derive an inferred attribute associated with the data subset;   associating the data subset with annotative data identifying the inferred attribute; and   converting the dataset in the first data format to an atomized dataset using another data format.

Join the waitlist — get patent alerts

Track US2023109821A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.