Management of artificial intelligence resources in a distributed resource environment
Abstract
Approaches presented herein provide for the management of artificial intelligence (AI)-related resources in a distributed resource environment, such as may be used to support accelerated machine learning (ML) applications on behalf of different users. Management functionality can be provided using an AI manager, such as a management service, that can determine the requirements, capabilities, and limitations of various available AI-related components, such as those of a plurality of AI models, engines, and accelerators, as well as the hardware (e.g., graphics processing units (GPUs)) that run or make up these AI-related resources. An AI manager can determine a selection and configuration of resources that is not only appropriate for use with a specific AI model, but that can also be optimized for factors such as throughput, resource utilization, and inference latency. An AI manager can ensure compatibility of resources and configuration, and can enforce access control to models and data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
receiving a request to perform inferencing using a specified machine learning model; determining a set of requirements associated with the specified machine learning model; identifying, in a distributed computing environment, a plurality of resources available to perform the inferencing; selecting, based at least on the set of requirements and resource compatibility information, a combination of the available resources of the distributed computing environment to perform the inferencing; performing the inferencing using the specified machine learning model with the selected combination of the available resources of the distributed computing environment; and providing one or more results of the inferencing in response to the request.
2 . The method of claim 1 , wherein the combination of the available resources includes at least a machine learning engine and a machine learning accelerator.
3 . The method of claim 1 , further comprising:
monitoring one or more available resources of the distributed computing environment to determine whether the one or more available resources are active and available to perform the inferencing before causing one or more new resource instances to be added to the plurality of resources.
4 . The method of claim 1 , wherein the request is received from an application executing outside of the distributed computing environment in which the plurality of available resources is located.
5 . The method of claim 4 , wherein the distributed computing environment is a multi-tenant environment in which at least a subset of the plurality of resources is able to be utilized by multiple users to perform different types of tasks.
6 . The method of claim 1 , further comprising:
selecting the combination of the available resources based further upon at least resource configuration information or resource version information.
7 . The method of claim 1 , further comprising:
selecting the combination of the available resources based at least further upon one or more access control policies corresponding to the request, the specified machine learning model, or the available resources.
8 . The method of claim 1 , further comprising:
releasing the selected combination of the available resources after performing the inferencing.
9 . A system, comprising:
one or more processing units to:
receive a request to perform an operation using at least one artificial intelligence model;
identify, in a multi-tenant resource environment, a plurality of available resources available to perform the operation;
determine, based at least on one or more requirements for the operation and resource compatibility information, a combination of the available resources to perform the inferencing; and
provide the selected combination of the available resources to perform the operation using the at least one artificial intelligence model.
10 . The system of claim 9 , wherein the one or more processing units are further to:
monitor one or more available resources to determine whether the one or more available resources are active and available to perform the inferencing before causing one or more new resource instances to be added to the plurality of resources.
11 . The system of claim 9 , wherein the combination of the available resources includes at least a machine learning engine and a machine learning accelerator.
12 . The system of claim 9 , wherein the one or more processing units are further to:
select the combination of the available resources based further upon at least resource configuration information or resource version information.
13 . The system of claim 9 , wherein the one or more processing units are further to:
select the combination of the available resources based further upon one or more access control policies corresponding to the request, the specified machine learning model, or the available resources.
14 . The system of claim 9 , wherein the system comprises at least one of:
a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
15 . The system of claim 9 , wherein the one or more processing units are further to:
release the selected combination of the available resources after performing the operation.
16 . A processor, comprising:
one or more circuits to:
receive a request to perform an operation using at least one artificial intelligence model;
identify, in a multi-tenant resource environment, a plurality of available resources available to perform the operation;
select, based at least on one or more requirements for the operation and resource compatibility information, a combination of the available resources to perform the inferencing; and
provide the selected combination of the available resources to perform the operation using the at least one artificial intelligence model.
17 . The processor of claim 16 , wherein the one or more circuits are further to:
monitor one or more available resources to determine whether the one or more available resources are active and available to perform the inferencing before causing one or more new resource instances to be added to the plurality of resources.
18 . The processor of claim 16 , wherein the combination of the available resources includes at least a machine learning engine and a machine learning accelerator.
19 . The processor of claim 16 , wherein the one or more circuits are further to:
select the combination of the available resources based further upon at least resource configuration information or resource version information.
20 . The processor of claim 16 , wherein the one or more circuits are further to:
select the combination of the available resources based further upon one or more access control policies corresponding to the request, the specified machine learning model, or the available resources.Join the waitlist — get patent alerts
Track US2024220831A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.