Scheduling system for computational work on heterogeneous hardware
Abstract
The technology includes methods, processes, and systems for virtualizing graphics processing unit (GPU) memory. Example embodiments of the technology include managing an amount of GPU memory used by one or more processes, such as Application Programming Interfaces (APIs), that directly or indirectly impact one or more other processes running on the same GPU. Managing and/or virtualizing the amount of GPU memory may ensure that an end user does not receive a GPU out-of-memory error because the API request is impacted by the processing of other API requests. A virtual machine with access to a GPU may be organized with one or more job slots that are configured to specify the number of processes that are able to run concurrently on a specific virtual machine. A process may be configured on each virtual machine running a software program or API and is used to schedule work based on GPU memory requirements.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
under the control of one or more computer systems configured with executable instructions,
managing an amount of Graphics Processing Unit (GPU) memory used by one or more processes, wherein the one or more processes directly or indirectly impact one or more other processes running on the GPU;
organizing a host machine with access to the GPU according to one or more request slots configured to specify a number of processes that are available to be processed by the GPU; and
scheduling the one or more processes based at least in part on the GPU memory.
2 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes Application Programming Interface (API) requests, complex-interaction calculations, neural networks, artificial intelligence, or other computation-intensive application.
3 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes:
optimizing a queue, in response to receiving two or more requests at or around a same time; determining an amount of time to process a first request of the two or more requests; determining an amount of time to process a second request of the two or more requests; and ordering the first request and the second request in a queue, wherein the queue is configured to store the two or more requests.
4 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes:
receiving, from a client device, a request to add a persistent slot; scheduling one process of the one or more processes in the persistent slot; and determining if the one process executes one or more child processes.
5 . A system, comprising:
at least one computing device configured to implement one or more services, wherein the one or more services are configured to:
manage an amount of Graphics Processing Unit (GPU) memory used by one or more processes, wherein the one or more processes directly or indirectly impact one or more other processes running on the GPU;
organize a host machine with access to the GPU according to one or more request slots configured to specify a number of processes that are available to be processed by the GPU; and
to schedule the one or more processes based at least in part on the GPU memory.
6 . The system of claim 5 , wherein the one or more processes are received from one or more client devices.
7 . The system of claim 5 , wherein the at least one computing device is further configured to:
receive, from a client device, a request to add a persistent slot; schedule one process of the one or more processes in the persistent slot; and determine if the one process executes one or more child processes.
8 . The system of claim 5 , wherein the at least one computing device is further configured to:
optimize a queue, in response to receiving two or more requests at or around a same time; determine an amount of time to process a first request of the two or more requests; determine an amount of time to process a second request of the two or more requests; and order the first request and the second request in a queue, wherein the queue is configured to store the two or more requests.
9 . The system of claim 8 , wherein the at least one computing device is further configured to:
receive at least one of information, input, and data associated with the one or more processes; and store the at least one of information, input, and data in a database operably connected to the queue.
10 . The system of claim 5 , wherein the host machine is a virtual machine operably connected to a GPU.
11 . The system of claim 5 , wherein one or more processes include Application Programming Interface (API) requests, complex-interaction calculations, neural networks, artificial intelligence, or other computation-intensive applications.
12 . The system of claim 5 , wherein the at least one computing device is further configured to:
launch a new host machine; create a new slot on an existing host machine; and designate additional resources to the host machine from a pool of host machines operably connected to the GPU.
13 . A non-transitory computer-readable storage medium having stored thereon executable instructions that, when executed by one or more processors of a computer system, cause the computer system to at least:
receive, from a client device, a request to process data associated with the request; schedule the request to one or more resources of a Graphics Processing Unit (GPU); identify an amount of GPU resources being available to process the request using at least the data associated with the request; and assign the request to the GPU.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to provide, to the client device, the processed request.
15 . The non-transitory computer-readable storage medium of claim 14 , wherein the instructions that cause the computer system to provide the processed request further include instructions that cause the computer system to maintain the request assigned to the GPU after the processed request is provided to the client device.
16 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to maintain at least one of the request, the data associated with the request, and the processed request.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to:
receive one or more additional requests, the one or more additional requests including new data associated with the request; retrieve the request maintained by the computer system; and process the one or more additional requests based on the request.
18 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions that cause the computer system to schedule the request further include instructions that, when executed by the one or more processors, cause the computer system to schedule the request in a slot of the GPU, wherein the slot of the GPU is initialized based on the amount of GPU resources being available.
19 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions that cause the computer system to identify the amount of GPU resources being available to process the request further include instructions that, when executed by the one or more processors, cause the computer system to determine a status of one or more slots of the GPU, wherein the one or more slots of the GPU are configured to receive at least the request and the data associated with the request in order to process the request.
20 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to:
determine if one or more other requests are assigned to the GPU; identify the one or more other requests assigned to the GPU to be evicted from the GPU; and evict the one or more other requests from the GPU.Join the waitlist — get patent alerts
Track US2022083395A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.