Memory interleaving coordinated by networked processing units
Abstract
Various approaches for configuring interleaving in a memory pool used in an edge computing arrangement, including with the use of infrastructure processing units (IPUs) and similar networked processing units, are disclosed. An example system may discover and map disaggregated memory resources at respective compute locations connected to each another via at least one interconnect. The system may identify workload requirements for use of the compute locations by respective workloads, for workloads provided by client devices to the compute locations. The system may determine an interleaving arrangement for a memory pool that fulfills the workload requirements, to use the interleaving arrangement to distribute data for the respective workloads among the disaggregated memory resources. The system may configure the memory pool for use by the client devices of the network, as the memory pool causes the disaggregated memory resources to host data based on the interleaving arrangement.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for configuring interleaving in a memory pool established in an edge computing arrangement, comprising:
discovering disaggregated memory resources at respective compute locations, the compute locations connected to each another via at least one interconnect; identifying workload requirements for use of the compute locations by respective workloads, the workloads provided by client devices to the compute locations via a network; determining an interleaving arrangement for a memory pool that fulfills the workload requirements, the interleaving arrangement to distribute data for the respective workloads among the disaggregated memory resources at the respective compute locations; and configuring the memory pool for use by the client devices of the network, the memory pool to cause the disaggregated memory resources among the compute locations to host data based on the interleaving arrangement.
2 . The method of claim 1 , wherein the method is performed by a network switch, and wherein the method further comprises:
processing requests, at the network switch, for the use of the memory pool by the client devices of the network.
3 . The method of claim 1 , wherein the method is performed by a networked processing unit, and wherein the method further comprises:
implementing, at the networked processing unit, the interleaving arrangement among the disaggregated memory resources by configuration of respective networked processing units at the respective compute locations.
4 . The method of claim 1 , wherein the workload requirements are identified based on one or more of:
a latency measurement for use of compute resources at the respective compute locations; an estimation of an availability of acceleration resources for current workloads in the network; a prediction of an availability of acceleration resources for future workloads in the network; a latency measurement for communications in the network; an estimation of current traffic in the network; or a prediction of bandwidth or load requirements in the network.
5 . The method of claim 1 , wherein the respective compute locations correspond to processing hardware at respective base stations, and wherein the client devices connect to the network via one or more of the respective base stations.
6 . The method of claim 1 , wherein one or more of the respective compute locations include acceleration resources, and wherein the disaggregated memory resources are mapped to the acceleration resources.
7 . The method of claim 6 , wherein the disaggregated memory resources are connected to the acceleration resources via a Compute Express Link (CXL) interconnect.
8 . The method of claim 1 , further comprising:
categorizing memory bandwidth available at the disaggregated memory resources into multiple categories; wherein the interleaving arrangement is determined using the multiple categories.
9 . The method of claim 1 , further comprising:
allocating a portion of the disaggregated memory resources at one or more compute locations without interleaving based on the workload requirements.
10 . The method of claim 1 , further comprising:
storing data in the memory pool according to the interleaving arrangement; and retrieving data in the memory pool according to the interleaving arrangement.
11 . The method of claim 1 , further comprising:
determining an updated interleaving arrangement; and reconfiguring the memory pool for use by the client devices, based on the updated interleaving arrangement.
12 . A device, comprising:
a networked processing unit; and a storage medium including instructions embodied thereon, wherein the instructions, which when executed by the networked processing unit, configure the networked processing unit to:
discover disaggregated memory resources at respective compute locations, the compute locations connected to each another via at least one interconnect;
identify workload requirements for use of the compute locations by respective workloads, the workloads provided by client devices to the compute locations via a network;
determine an interleaving arrangement for a memory pool that fulfills the workload requirements, the interleaving arrangement to distribute data for the respective workloads among the disaggregated memory resources at the respective compute locations; and
configure the memory pool for use by the client devices of the network, the memory pool to cause the disaggregated memory resources among the compute locations to host data based on the interleaving arrangement.
13 . The device of claim 12 , wherein the device is a network switch, and wherein the instructions further configure the networked processing unit to:
process requests, at the network switch, for the use of the memory pool by the client devices of the network.
14 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
provide commands to respective networked processing units at the respective compute locations, to cause the respective networked processing units to implement the interleaving arrangement among the disaggregated memory resources.
15 . The device of claim 12 , wherein the workload requirements are identified based on one or more of:
a latency measurement for use of compute resources at the respective compute locations; an estimation of an availability of acceleration resources for current workloads in the network; a prediction of an availability of acceleration resources for future workloads in the network; a latency measurement for communications in the network; an estimation of current traffic in the network; or a prediction of bandwidth or load requirements in the network.
16 . The device of claim 12 , wherein the respective compute locations correspond to processing hardware at respective base stations, and wherein the client devices connect to the network via one or more of the respective base stations.
17 . The device of claim 12 , wherein one or more of the respective compute locations include acceleration resources, and wherein the disaggregated memory resources are mapped to the acceleration resources.
18 . The device of claim 17 , wherein the disaggregated memory resources are connected to the acceleration resources via a Compute Express Link (CXL) interconnect.
19 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
categorize memory bandwidth available at the disaggregated memory resources into multiple categories; wherein the interleaving arrangement is determined using the multiple categories.
20 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
allocate a portion of the disaggregated memory resources at one or more compute locations without interleaving based on the workload requirements.
21 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
store data in the memory pool according to the interleaving arrangement; and retrieve data in the memory pool according to the interleaving arrangement.
22 . The device of claim 12 , wherein the instructions further configure the networked processing unit to:
determine an updated interleaving arrangement; and reconfigure the memory pool for use by the client devices, based on the updated interleaving arrangement.
23 . A non-transitory machine-readable storage medium comprising information representative of instructions, wherein the instructions, when executed by processing circuitry, cause the processing circuitry to:
select disaggregated memory resources from respective compute locations, the compute locations connected to each another via at least one interconnect; generate workload requirements for use of the compute locations by respective workloads, the workloads provided by client devices to the compute locations via a network; determine an interleaving arrangement for a memory pool that fulfills the workload requirements, the interleaving arrangement to distribute data for the respective workloads among the disaggregated memory resources at the respective compute locations; and configure the memory pool for use by the client devices of the network, the memory pool to cause the disaggregated memory resources among the compute locations to host data based on the interleaving arrangement.
24 . The non-transitory machine-readable storage medium of claim 23 , wherein the workload requirements are identified based on one or more of:
a latency measurement for use of compute resources at the respective compute locations; an estimation of an availability of acceleration resources for current workloads in the network; a prediction of an availability of acceleration resources for future workloads in the network; a latency measurement for communications in the network; an estimation of current traffic in the network; or a prediction of bandwidth or load requirements in the network.
25 . The non-transitory machine-readable storage medium of claim 23 , wherein the instructions further configure the processing circuitry to:
categorize memory bandwidth available at the disaggregated memory resources into multiple categories, wherein the interleaving arrangement is determined using the multiple categories; and allocate a portion of the disaggregated memory resources at one or more compute locations without interleaving based on the multiple categories.Join the waitlist — get patent alerts
Track US2023134683A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.