US2025037006A1PendingUtilityA1
Instance recommendations for machine learning workloads
Est. expiryJul 25, 2043(~17 yrs left)· nominal 20-yr term from priority
G06N 20/00
60
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
In various examples, a ranking is generated for a set of computing instances based on predicted metrics associated with computing instances. For example, a prediction model estimates various system performance metrics based on information associated with a workload and configuration information associated with computing instances. The system performance metrics estimated by the prediction model are used to rank the set of computing instances.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
obtaining an indication of a metric to rank a set of computing instances and a set of workload features of a workload; causing a machine learning model to determine an epoch training time and a processor utilization for computing instances of the set of computing instances based on the set of workload features of the workload, a set of computing instance features of the set of computing instances, and a set of performance features; ranking the set of computing instances in accordance with the metric based on the epoch training time and the processor utilization associated with the computing instances of the set of computing instances; and causing the ranking of the set of computing instances in a user interface.
2 . The method of claim 1 , wherein the method further comprises causing a second machine learning model to determine the set of performance features based on the set of workload features and the set of computing instance features.
3 . The method of claim 2 , wherein causing the second machine learning model to determine the set of performance features is in response to the workload or the set of computing instances having not been previously recorded.
4 . The method of claim 2 , wherein the method further comprises training the machine learning model and the second machine learning model using a training dataset including a set of metrics obtained by at least causing the set of computing instances to execute a set of workloads.
5 . The method of claim 1 , wherein the set of computing instance features includes at least one of: a number of Graphic Processing Units (GPUs), GPU memory, GPU memory type, GPU type, number of Central Processing Units (CPUs), number of virtual CPUs, CPU type, CPU memory, and CPU memory type.
6 . The method of claim 1 , wherein the set of performance features includes at least one of: average Graphic Processing Unit (GPU) utilization, minimum GPU utilization, maximum GPU utilization, average Central Processing Unit (CPU) utilization, minimum CPU utilization, maximum CPU utilization, average memory utilization, minimum memory utilization, maximum memory utilization, core temperature, memory bandwidth, cache usage, and power usage.
7 . The method of claim 1 , wherein the set of workload features includes at least one of: a number of floating point operations (FLOPs), number of layers, number of activations, number of parameters, and batch size.
8 . A non-transitory computer-readable medium storing executable instructions embodied thereon, which, when executed by a processing device, cause the processing device to perform operations comprising:
causing a machine learning model to determine a metric associated with a computing instance when executing a workload, the machine learning model taking as inputs a set of workload features associated with the workload, a set of computing instance features associated with a plurality of computing instances, and a set of metrics obtained from the plurality of computing instances during execution of a plurality of workloads; generating a ranking of a set of computing instances, including the computing instance, based on the metric; and updating a display to include the ranking of the set of computing instances.
9 . The medium of claim 8 , wherein the processing device further performs operations comprising causing a second machine learning model to determine at least a portion of the set of metrics.
10 . The medium of claim 8 , wherein the set of metrics include benchmarks obtained from the plurality of computing instances during execution of the plurality of workloads.
11 . The medium of claim 8 , wherein the machine learning model is trained using the set of metrics.
12 . The medium of claim 8 , wherein the ranking of the set of computing instances further comprises an ordering of the set of computing instances from a lowest epoch training time to a highest epoch training time.
13 . The medium of claim 8 , wherein the ranking of the set of computing instances further comprises an ordering of the set of computing instances from a highest processor utilization to a lowest processor utilization.
14 . The medium of claim 8 , wherein the machine learning model is a regression model.
15 . The medium of claim 8 , wherein the processing device further performs operations comprising:
obtaining an indication of the metric to optimize and the workload from a user interface; and determining the set of workload features based on the workload.
16 . A system comprising:
a memory component; and a processing device coupled to the memory component, the processing device to perform operations comprising:
obtaining a training dataset including benchmark data captured from a plurality of computing instance configurations executing a plurality of workloads;
training a machine learning model to determine system performance features for a set of computing instances based on a set of workload features and the set of computing instances, the machine learning model trained using the training dataset including a set of computing instance features and a set of machine learning model features extracted from the training dataset;
providing the machine learning model to an instance recommendation tool to rank computing instances based on the system performance features; and
causing the instance recommendation tool to rank computing instances based on the system performance features.
17 . The system of claim 16 , wherein the system performance features include at least one of: average Graphic Processing Unit (GPU) utilization, minimum GPU utilization, maximum GPU utilization, average Central Processing Unit (CPU) utilization, minimum CPU utilization, maximum CPU utilization, average memory utilization, minimum memory utilization, maximum memory utilization, core temperature, memory bandwidth, cache usage, and power usage.
18 . The system of claim 16 , wherein the set of computing instance features includes at least one of: a number of Graphic Processing Units (GPUs), GPU memory, GPU memory type, GPU type, number of Central Processing Units (CPUs), number of virtual CPUs, CPU type, CPU memory, and CPU memory type.
19 . The system of claim 16 , wherein the set of machine learning model features includes at least one of: a number of floating point operations (FLOPs), number of layers, number of activations, number of parameters, and batch size.
20 . The system of claim 16 , wherein the training dataset is generated by at least causing a set of machine learning models corresponding to the set of machine learning model features to executed the plurality of workloads using a plurality of computing instances.Join the waitlist — get patent alerts
Track US2025037006A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.