Management of workload processing using distributed networked processing units
Abstract
Various approaches for deploying and controlling distributed compute operations with the use of infrastructure processing units (IPUs) and similar networked processing units are disclosed. A system that includes a networked processing unit may perform workload processing with operations that: receive, from another networked processing unit, workload information for a workload, for a workload having respective tasks to be processed among distributed computing entities; perform an analysis of network conditions for a predicted execution of the workload, based on the workload information, to analyze network availability among the distributed computing entities; perform an analysis of compute conditions for the predicted execution of the workload, based on the workload information, to analyze processing availability among the distributed computing entities; and identify locations of the distributed computing entities to deploy the workload, based on the analysis of network conditions and the analysis of compute conditions.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for workload processing in an edge computing arrangement, comprising:
receiving, from a networked processing unit, workload information for a workload, the workload including respective tasks to be processed among distributed computing entities; performing an analysis of network conditions for a predicted execution of the workload, based on the workload information, to analyze network availability among the distributed computing entities; performing an analysis of compute conditions for the predicted execution of the workload, based on the workload information, to analyze processing availability among the distributed computing entities; and identifying locations of the distributed computing entities to deploy the workload, based on the analysis of network conditions and the analysis of compute conditions.
2 . The method of claim 1 , further comprising:
causing activation of the identified locations of the distributed computing entities to enable execution of the workload.
3 . The method of claim 1 , wherein the network availability relates to a measurement of congestion, priority, and latency of network connections used among the distributed computing entities.
4 . The method of claim 1 , wherein the processing availability relates to a measurement of compute resources located among the distributed computing entities.
5 . The method of claim 4 , wherein the compute resources include: at least a first plurality of central processing unit (CPU) cores located at a first computing node and at least a second plurality of CPU cores located at a second computing node.
6 . The method of claim 4 , wherein the compute resources include: at least a first plurality of accelerators located at a first computing node and at least a second plurality of accelerators located at a second computing node.
7 . The method of claim 1 , further comprising:
performing an analysis of power usage and power availability among the distributed computing entities; wherein identifying the locations in the distributed computing entities to deploy the workload is further based on the analysis of power usage and power availability.
8 . The method of claim 7 , further comprising:
causing a change to at least one power state at the identified locations of the distributed computing entities during deployment of the workload.
9 . The method of claim 1 , wherein the analysis of network conditions and the analysis of compute conditions each include evaluation of at least one priority of the workload and evaluation of at least one characteristic of a service level agreement for the workload.
10 . The method of claim 1 , wherein the method is performed by a networked processing unit in a network interface of a computing system, and wherein the workload information is received via the network interface from another networked processing unit located at a base station, on-premises server, or data center server.
11 . The method of claim 10 , further comprising:
providing commands to cause a network switch or gateway to control network traffic to the identified locations of the distributed computing entities during deployment of the workload.
12 . A device, comprising:
a networked processing unit; and a storage medium including instructions embodied thereon, wherein the instructions, which when executed by the networked processing unit, configure the networked processing unit to:
receive, from another networked processing unit, workload information for a workload, the workload including respective tasks to be processed among distributed computing entities;
perform an analysis of network conditions for a predicted execution of the workload, based on the workload information, to analyze network availability among the distributed computing entities;
perform an analysis of compute conditions for the predicted execution of the workload, based on the workload information, to analyze processing availability among the distributed computing entities; and
identify locations of the distributed computing entities to deploy the workload, based on the analysis of network conditions and the analysis of compute conditions.
13 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
cause activation of the identified locations of the distributed computing entities to enable execution of the workload.
14 . The device of claim 12 , wherein the network availability relates to a measurement of congestion, priority, and latency of network connections used among the distributed computing entities.
15 . The device of claim 12 , wherein the processing availability relates to a measurement of compute resources located among the distributed computing entities.
16 . The device of claim 15 , wherein the compute resources include: at least a first plurality of central processing unit (CPU) cores located at a first computing node and at least a second plurality of CPU cores located at a second computing node.
17 . The device of claim 15 , wherein the compute resources include: at least a first plurality of accelerators located at a first computing node and at least a second plurality of accelerators located at a second computing node.
18 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
perform an analysis of power usage and power availability among the distributed computing entities; wherein identifying the locations in the distributed computing entities to deploy the workload is further based on the analysis of power usage and power availability.
19 . The device of claim 18 , wherein the instructions further configure the networked processing unit to:
cause a change to at least one power state at the identified locations of the distributed computing entities during deployment of the workload.
20 . The device of claim 12 , wherein the analysis of network conditions and the analysis of compute conditions each include evaluation of at least one priority of the workload and evaluation of at least one characteristic of a service level agreement for the workload.
21 . The device of claim 12 , wherein the networked processing unit is provided in a network interface of the device, and wherein the workload information is received via the network interface from another networked processing unit located at a base station, on-premises server, or data center server.
22 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
provide commands to cause a network switch or gateway to control network traffic to the identified locations of the distributed computing entities during deployment of the workload.
23 . A non-transitory machine-readable storage medium comprising information representative of instructions, wherein the instructions, when executed by a networked processing unit, cause the networked processing unit to:
receive, from another networked processing unit, workload information for a workload, the workload including respective tasks to be processed among distributed computing entities; perform an analysis of network conditions for a predicted execution of the workload, based on the workload information, to analyze network availability among the distributed computing entities; perform an analysis of compute conditions for the predicted execution of the workload, based on the workload information, to analyze processing availability among the distributed computing entities; and identify locations of the distributed computing entities to deploy the workload, based on the analysis of network conditions and the analysis of compute conditions.
24 . The non-transitory machine-readable storage medium of claim 23 , wherein the instructions further configure the networked processing unit to:
cause activation of the identified locations of the distributed computing entities to enable execution of the workload; wherein the network availability relates to a measurement of congestion, priority, and latency of network connections used among the distributed computing entities; and wherein the processing availability relates to a measurement of compute resources located among the distributed computing entities.
25 . The non-transitory machine-readable storage medium of claim 24 , wherein the compute resources include:
at least a first plurality of central processing unit (CPU) cores located at a first computing node and at least a second plurality of CPU cores located at a second computing node; or at least a first plurality of accelerators located at the first computing node and at least a second plurality of accelerators located at the second computing node.Join the waitlist — get patent alerts
Track US2023135645A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.