US2023164219A1PendingUtilityA1

Access Pattern Driven Data Placement in Cloud Storage

Assignee: GOOGLE LLCPriority: Nov 4, 2019Filed: Jan 19, 2023Published: May 25, 2023
Est. expiryNov 4, 2039(~13.3 yrs left)· nominal 20-yr term from priority
H04L 67/06H04L 67/52G06F 3/0647H04L 67/568G06F 3/0611G06N 5/01G06N 20/00H04L 67/5681H04L 67/1097G06F 16/172G06F 3/067H04L 67/535G06F 16/1824G06F 16/183H04L 65/4015
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for storing data in a distributed network having a plurality of datacenters distributed over a plurality of geographic regions. The method may involve receiving data, including metadata, uploaded to a first datacenter of the distributed network, receiving access information about previous data that was previously stored in the plurality of datacenters of the distributed network, predicting one or more of the plurality of geographic regions from which the uploaded data will be accessed based on the metadata and the access information, and instructing the uploaded data to be transferred from the first datacenter to one or more second datacenters located at each of the one or more predicted geographic regions.

Claims

exact text as granted — not AI-modified
1 . A method for storing a plurality of data items in a distributed network having a plurality of datacenters distributed over a plurality of geographic regions, the method comprising:
 receiving, by one or more processors, a plurality of first data items uploaded to the distributed network from a plurality of first users, each first data item including metadata, the metadata including an upload geographic region at which the first data item is uploaded and one or more accessed geographic regions at which the first data item is accessed;   training, by the one or more processors, a predictive model using the metadata of the plurality of first data items;   after training the predictive model using the metadata of the plurality of first data items, receiving, by the one or more processors, a second data item uploaded to the distributed network by a second user;   determining, by the one or more processors, one or more storage geographic regions at which the second data item is to be stored based at least in part on the predictive model, wherein at least one of the one or more storage geographic locations at which the second data item is to be stored is different from the upload geographic region at which the second data item was uploaded; and   instructing, by the one or more processors, the second data item to be transferred from the upload geographic region of the second data item to one or more datacenters of the one or more storage geographic regions at which the second data item is to be stored.   
     
     
         2 . The method of  claim 1 , further comprising predicting, by the one or more processors, one or more access geographic regions at which the second data item is predicted to be accessed based on the predictive model, wherein the one or more storage geographic regions at which the second data item is to be stored are determined based on the predicted one or more access geographic regions. 
     
     
         3 . The method of  claim 1 , wherein the predictive model is a decision tree model. 
     
     
         4 . The method of  claim 1 , wherein the metadata further includes, and the predictive model is trained with, at least one of:
 an identification of a datacenter to which the first data item is uploaded;   an identification of an uploading user;   a time of upload;   a size of the first data item; or   a name of the first data item.   
     
     
         5 . The method of  claim 1 , wherein the metadata further includes, and the predictive model is trained with, at least one of:
 one or more second storage geographic regions at which the first data item is stored; or   one or more times at which the first data item is accessed; or   a number of access requests for the first data item.   
     
     
         6 . The method of  claim 1 , wherein the plurality of first data items includes at least one data file, wherein the metadata of the data file includes file characteristic data, wherein the file characteristic data includes at least one of: a name of the file; a size of the file; or an identification of a directory or a file path at which the file is stored, and wherein the predictive model is trained at least in part using the file characteristic data. 
     
     
         7 . The method of  claim 1 , further comprising:
 predicting, by the one or more processors, an amount of time until the second data item will be accessed for a first time;   for at least one of the determined storage geographic regions of the second data item, selecting, by the one or more processors, one of a first transfer protocol or a second transfer protocol for transferring the second data item to the at least one storage geographic region, based on the predicted amount of time, wherein an average time for the second data item to arrive at the at least one storage geographic region using the first transfer protocol is less than an average time for the second data item to arrive at the at least one storage geographic region using the second transfer protocol; and   transferring, by the one or more processors, the second data item from the upload geographic region of the second data item to the at least one storage geographic region of the second data item according to the selected first or second transfer protocol.   
     
     
         8 . The method of  claim 7 , wherein the first transfer protocol comprises cache injection of the second data item to one or more caching servers located at the at least one storage geographic region. 
     
     
         9 . The method of  claim 8 , wherein the second transfer protocol comprises:
 instructing, by the one or more processors, the second data item to be included in a file including other uploaded data items having a common storage geographic region as the second data item; and   instructing, by the one or more processors, the file to be transferred to one or more datacenters located at the common storage geographic region.   
     
     
         10 . The method of  claim 9 , wherein the first transfer protocol comprises transferring the second data item according to the second transfer protocol in addition to the cache injection. 
     
     
         11 . A system for storing a plurality of data items in a distributed network having a plurality of datacenters distributed over a plurality of geographic regions, the system comprising:
 one or more storage devices configured to store a plurality of first data items uploaded to the distributed network from a plurality of first users, each first data item including metadata, the metadata including an upload geographic region at which the first data item is uploaded and one or more accessed geographic regions at which the first data item is accessed; and   one or more processors in communication with the one or more storage devices, the one or more processors configured to:
 train a predictive model using the metadata of the plurality of first data items; 
 after training the predictive model using the metadata of the plurality of first data items, for a second data item uploaded to the distributed network by a second user:
 determine one or more storage geographic regions at which the second data item is to be stored based at least in part on the predictive model, wherein at least one of the one or more storage geographic locations at which the second data item is to be stored is different from the upload geographic region at which the second data item was uploaded; and 
 
 instruct the second data item to be transferred from the upload geographic region of the second data item to one or more datacenters of the one or more storage geographic regions at which the second data item is to be stored. 
   
     
     
         12 . The system of  claim 11 , wherein the one or more processors are configured to predict one or more access geographic regions at which the second data item is predicted to be accessed based on the predictive model, wherein the one or more storage geographic regions at which the second data item is to be stored are determined based on the predicted one or more access geographic regions. 
     
     
         13 . The system of  claim 11 , wherein the predictive model is a decision tree model. 
     
     
         14 . The system of  claim 11 , wherein the metadata further includes, and the predictive model is trained on, at least one of:
 an identification of a datacenter to which the first data item is uploaded;   an identification of an uploading user;   a time of upload;   a size of the first data item; or   a name of the first data item.   
     
     
         15 . The system of  claim 1 , wherein the metadata further includes, and the predictive model is trained on, at least one of:
 one or more second storage geographic regions at which the first data item is stored; or   one or more times at which the first data item is accessed; or   a number of access requests for the first data item.   
     
     
         16 . The system of  claim 14 , wherein the plurality of first data items includes at least one data file, wherein the metadata of the data file includes file characteristic data, wherein the file characteristic data includes at least one of: a name of the file; a size of the file; or an identification of a directory or a file path at which the file is stored, and wherein the one or more processors are configured to train the predictive model based at least in part on the file characteristic data. 
     
     
         17 . The system of  claim 11 , wherein the one or more processors are configured to:
 predict an amount of time until the second data item will be accessed for a first time; and   for at least one of the determined storage geographic regions of the second data item, select one of a first transfer protocol or a second transfer protocol for transferring the second data item to the at least one storage geographic region, based on the predicted amount of time, wherein an average time for the second data item to arrive at the at least one storage geographic region using the first transfer protocol is less than an average time for the second data item to arrive at the at least one storage geographic region using the second transfer protocol; and   transfer the second data item from the upload geographic region of the second data item to the at least one storage geographic region of the second data item according to the selected first or second transfer protocol.   
     
     
         18 . The system of  claim 17 , wherein the first transfer protocol comprises cache injection of the second data item to one or more caching servers located at the at least one storage geographic region. 
     
     
         19 . The system of  claim 11 , wherein the second transfer protocol comprises:
 instruction of the second data item to be included in a file including other uploaded data items having a common storage geographic region as the second data item; and   instruction of the file to be transferred to one or more datacenters located at the common storage geographic region.   
     
     
         20 . The system of  claim 19 , wherein the first transfer protocol comprises performance of the second transfer protocol in addition to the cache injection.

Join the waitlist — get patent alerts

Track US2023164219A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.