Ingestion partition auto-scaling in a time-series database
Abstract
Methods, systems, and computer-readable media for ingestion partition auto-scaling in a time-series database are disclosed. A first set of one or more hosts divides elements of time-series data into a plurality of partitions. A second set of one or more hosts stores the elements of time-series data from the plurality of partitions into one or more storage tiers of a time-series database. An analyzer receives first data indicative of the resource usage of the time-series data at the first set of one or more hosts. The analyzer receives second data indicative of the resource usage of the time-series data at the second set of one or more hosts. Based at least in part on analysis of the first data and the second data, the analyzer initiates a split of an individual one of the partitions into two or more partitions.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A system, comprising:
one or more hosts, comprising one or more processors and associated memory, configured to:
store the elements of time-series data into two-dimensional tiles at at least one storage tier of a time-series database, wherein the two-dimensional tiles are divided into spatial partitions along a first dimension, and divided into temporal partitions along a second dimension;
an analyzer, comprising one or more processors and associated memory, configured to:
receive data indicative of a resource usage of the time-series data at the one or more hosts that store the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database; and
initiate a split of an individual one of the spatial partitions into two or more spatial partitions, wherein the split is initiated based at least in part on an analysis of the data.
22 . The system as recited in claim 1 , further comprising:
a second set of one or more hosts, comprising one or more processors and associated memory, configured to:
divide the elements of the time-series data into a plurality of spatial partitions based at least in part on a spatial range;
wherein to store the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database, the one or more hosts are further configured to:
store the elements of the time-series data from the plurality of spatial partitions into the two-dimensional tiles at the at least one storage tier of the time-series database, wherein the two-dimensional tiles are divided into the plurality of spatial partitions along the first dimension, and divided into the temporal partitions along the second dimension; and
wherein the analyzer is further configured to:
receive second data indicative of a resource usage of the time-series data at the second set of one or more hosts that divides the elements of the time-series data into the plurality of spatial partitions;
wherein the split of the individual one of the spatial partitions into the two or more spatial partitions is initiated based at least in part on an analysis of the data and the second data.
23 . The system as recited in claim 21 , wherein the one or more hosts comprise one or more stream processors of the time-series database, and wherein the system further comprises:
one or more ingestion routers of the time-series database, comprising one or more processors and associated memory, configured to:
divide the elements of the time-series data into a plurality of ingestion partitions;
wherein a particular ingestion partition is assigned to one and only one stream processor of the one or more stream processors, for storing the elements of the time-series data from the particular ingestion partition into the two-dimensional tiles at the at least one storage tier of the time-series database.
24 . The system as recited in claim 1 , wherein the analyzer is further configured to:
based at least in part on the analysis of the data, initiate a merge of two or more of the spatial partitions into a single partition.
25 . The system as recited in claim 21 , wherein the spatial partitions are non-overlapping spatial partitions, wherein the temporal partitions are non-overlapping temporal partitions, and wherein the spatial partitions along the first dimension are divided based at least in part on a hierarchical clustering that co-locates related measurements or time-series into the same spatial partition.
26 . The system as recited in claim 21 , wherein the two-dimensional tiles comprise open tiles and closed tiles, wherein open tiles represent a current window of time, and closed tiles represent older windows of time, and wherein an open tile is closed when an amount of data of the open tile reaches a threshold, or when a maximum time interval for the open tile is reached.
27 . The system as recited in claim 21 , wherein the contents of at least some of the two-dimensional tiles are replicated to different storage tiers at different locations.
28 . A method, comprising:
storing, by a set of one or more hosts, the elements of time-series data into two-dimensional tiles at at least one storage tier of a time-series database, wherein the two-dimensional tiles are divided into spatial partitions along a first dimension, and divided into temporal partitions along a second dimension; receiving, by an analyzer, data indicative of a resource usage of the time-series data at the one or more hosts that store the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database; and initiating, by the analyzer, a split of an individual one of the spatial partitions into two or more spatial partitions, wherein the split is initiated based at least in part on an analysis of the data.
29 . The method as recited in claim 28 , further comprising:
dividing, by a second set of one or more hosts, the elements of the time-series data into a plurality of spatial partitions based at least in part on a spatial range; wherein the storing of the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database, further comprises:
storing the elements of the time-series data from the plurality of spatial partitions into the two-dimensional tiles at the at least one storage tier of the time-series database, wherein the two-dimensional tiles are divided into the plurality of spatial partitions along the first dimension, and divided into the temporal partitions along the second dimension; and
wherein the method further comprises:
receiving, at the analyzer, second data indicative of a resource usage of the time-series data at the second set of one or more hosts that divides the elements of the time-series data into the plurality of spatial partitions;
wherein the split of the individual one of the spatial partitions into the two or more spatial partitions is initiated based at least in part on an analysis of the data and the second data.
30 . The method as recited in claim 28 , wherein the one or more hosts comprise one or more stream processors of the time-series database, and wherein the method further comprises:
dividing, by one or more ingestion routers of the time-series database, the elements of the time-series data into a plurality of ingestion partitions; wherein a particular ingestion partition is assigned to one and only one stream processor of the one or more stream processors, for storing the elements of the time-series data from the particular ingestion partition into the two-dimensional tiles at the at least one storage tier of the time-series database.
31 . The method as recited in claim 28 , further comprising:
based at least in part on the analysis of the data, initiating by the analyzer a merge of two or more of the spatial partitions into a single partition.
32 . The method as recited in claim 28 , wherein the spatial partitions are non-overlapping spatial partitions, wherein the temporal partitions are non-overlapping temporal partitions, and wherein the spatial partitions along the first dimension are divided based at least in part on a hierarchical clustering that co-locates related measurements or time-series into the same spatial partition.
33 . The method as recited in claim 28 , wherein the two-dimensional tiles comprise open tiles and closed tiles, wherein open tiles represent a current window of time, and closed tiles represent older windows of time, and wherein an open tile is closed when an amount of data of the open tile reaches a threshold, or when a maximum time interval for the open tile is reached.
34 . The method as recited in claim 28 , wherein tiles of the two-dimensional tiles, whose temporal boundaries are beyond a retention period, are either deleted or marked for deletion.
35 . One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more processors, cause the one or more processors to:
store, by a set of one or more hosts, the elements of time-series data into two-dimensional tiles at at least one storage tier of a time-series database, wherein the two-dimensional tiles are divided into spatial partitions along a first dimension, and divided into temporal partitions along a second dimension; receive, by an analyzer, data indicative of a resource usage of the time-series data at the one or more hosts that store the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database; and initiate, by the analyzer, a split of an individual one of the spatial partitions into two or more spatial partitions, wherein the split is initiated based at least in part on an analysis of the data.
36 . The one or more non-transitory computer-readable storage media as recited in claim 35 , further comprising additional program instructions that, when executed on or across the one or more processors, perform:
dividing, by a second set of one or more hosts, the elements of the time-series data into a plurality of spatial partitions based at least in part on a spatial range; and receiving, at the analyzer, second data indicative of a resource usage of the time-series data at the second set of one or more hosts that divides the elements of the time-series data into the plurality of spatial partitions; wherein the split of the individual one of the spatial partitions into the two or more spatial partitions is initiated based at least in part on an analysis of the data and the second data; and wherein the storing of the elements of the time-series data into the two-dimensional tiles at the at least one storage tier of the time-series database, further comprises:
storing the elements of the time-series data from the plurality of spatial partitions into the two-dimensional tiles at the at least one storage tier of the time-series database, wherein the two-dimensional tiles are divided into the plurality of spatial partitions along the first dimension, and divided into the temporal partitions along the second dimension.
37 . The one or more non-transitory computer-readable storage media as recited in claim 36 , further comprising additional program instructions that, when executed on or across the one or more processors, perform:
determining, by the analyzer, a split point for the split in the individual one of the spatial partitions, wherein the split point is determined based at least in part on the second data.
38 . The one or more non-transitory computer-readable storage media as recited in claim 35 , wherein the data represents partition-specific heat data over a window of time, and wherein the data is pushed to the analyzer.
39 . The one or more non-transitory computer-readable storage media as recited in claim 35 , wherein the split is delayed based at least in part on a rate limit associated with partition splits in the time-series database.
40 . The one or more non-transitory computer-readable storage media as recited in claim 35 , further comprising additional program instructions that, when executed on or across the one or more processors, perform:
initiating, by the analyzer, a defragmentation of two or more of the spatial partitions, wherein the defragmentation comprises one or more splits and one or more merges, and wherein the defragmentation is initiated based at least in part on the analysis of the data.Join the waitlist — get patent alerts
Track US2022171792A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.