Column ordering for input/output optimization in tabular data
Abstract
Systems, methods, and computer-readable media for determining column ordering of a data storage table for search optimization are described herein. In some examples, a computing system is configured to receive input containing statistics of a plurality of queries. The computing system can then determine a new column order (i.e., layout) based at least in part on the statistics. In some example techniques described herein, the computing system can determine the new column order based at least in part on the hardware components storing the data storage table, storage system parameters, and/or user preference information. Example techniques described herein can apply the new column order to data subsequently added to the data storage table. Example techniques described herein can apply the new column order to existing data in the data storage table.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A system comprising:
a processor; memory storing instructions that, when executed by the processor, causes the system to perform operations comprising:
identifying two related columns of a plurality of columns stored in a data storage table, wherein the plurality of columns is arranged in a first order;
determining a first seek cost to access the two related columns;
determining the two related columns are not located close to each other based on the first seek cost;
in response to determining the two related columns are not located close to each other, determining a second order for the plurality of columns based on one or more storage system parameters associated with the data storage table; and
arranging the plurality of columns according to the second order.
22 . The system of claim 21 , wherein identifying two related columns comprises a determination that the two related columns are likely to be accessed together in a query.
23 . The system of claim 22 , wherein the determination that the two related columns are likely to be accessed together in the query is based on query statistics comprising at least one of:
a number of queries; a type of query; or a list of columns accessed per query.
24 . The system of claim 21 , wherein the first seek cost identifies an amount of time used to access a second column of the two related columns after accessing a first column of the two related columns.
25 . The system of claim 24 , wherein determining the two related columns are not located close to each other comprises determining that the amount of time used to access the second column after accessing the first column exceeds a threshold time value.
26 . The system of claim 24 , wherein the amount of time used to access the second column after accessing the first column is based on at least one of a track seek time or a disk rotation time.
27 . The system of claim 21 , wherein determining a second order for the plurality of columns comprises determining that a second seek cost for accessing the two related columns in the second order is less than the first seek cost for accessing the two related columns in the first order.
28 . The system of claim 21 , wherein arranging the plurality of columns according to the second order comprises arranging the two related columns such that the two related columns are located close together.
29 . The system of claim 21 , wherein arranging the two related columns such that the two related columns are located close together comprises storing the two related columns adjacent to one another in the data storage table.
30 . The system of claim 21 , wherein the one or more storage system parameters comprise at least one of:
a column compression; a row size; a data type; or a row type.
31 . The system of claim 21 , wherein arranging the plurality of columns according to the second order comprises:
determining that at least one of the one or more storage system parameters exceeds a threshold value; and causing the system to automatically arrange the plurality of columns based on exceeding the threshold value.
32 . The system of claim 21 , the operations further comprising:
subsequent to arranging the plurality of columns according to the second order, receiving a search query; and performing the search query in accordance with the second order.
33 . A method comprising:
identifying two related columns of a plurality of columns stored in a data storage table, wherein the plurality of columns is arranged in a first order; determining a first seek cost to access the two related columns; determining the two related columns are not located close to each other based on the first seek cost; in response to determining the two related columns are not located close to each other, determining a second order for the plurality of columns based on one or more storage system parameters associated with the data storage table; and arranging the plurality of columns according to the second order.
34 . The method of claim 33 , wherein determining the first seek cost comprises:
receiving one or more query statistics for the data storage table; and processing the one or more query statistics to determine the first seek cost.
35 . The method of claim 34 , wherein processing the one or more query statistics comprises determining seek costs for a set of queries.
36 . The method of claim 34 , wherein the second order is further determined based on the one or more query statistics.
37 . The method of claim 33 , the method further comprising:
receiving user preference data including instructions to enable or disable automatic column ordering of the plurality of columns; and arranging the plurality of columns according to the second order based on the user preference data.
38 . The method of claim 33 , the method further comprising:
determining a second seek cost to access the two related columns in the second order; and arranging the plurality of columns according to the second order when the second seek cost is less than the first seek cost.
39 . The method of claim 33 , wherein the two related columns are closer to each other in the second order than in the first order.
40 . A device comprising:
a processor; memory storing instructions that, when executed by the processor, causes the device to perform operations comprising:
identifying two related columns of a plurality of columns stored in a data storage table, wherein the plurality of columns is arranged in a first order;
determining a seek cost to access the two related columns;
determining the two related columns are not located close to each other based on the seek cost;
in response to determining the two related columns are not located close to each other, determining a second order for the plurality of columns based on one or more storage system parameters associated with the data storage table; and
arranging the plurality of columns according to the second order.Join the waitlist — get patent alerts
Track US2023078315A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.