Remote direct memory access (rdma)-based recovery of dirty data in remote memory
Abstract
Techniques for implementing RDMA-based recovery of dirty data in remote memory are provided. In one set of embodiments, upon occurrence of a failure at a first (i.e., source) host system, a second (i.e., failover) host system can allocate a new memory region corresponding to a memory region of the source host system and retrieve a baseline copy of the memory region from a storage backend shared by the source and failover host systems. The failover host system can further populate the new memory region with the baseline copy and retrieve one or more dirty page lists for the memory region from the source host system via RDMA, where the one or more dirty page lists identify memory pages in the memory region that include data updates not present in the baseline copy. For each memory page identified in the one or more dirty page lists, the failover host system can then copy the content of that memory page from the memory region of the source host system to the new memory region via RDMA.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising, upon occurrence of a failure at a source host system:
allocating, by a failover host system, a new memory region in a physical memory of the failover host system that corresponds to a memory region in a physical memory of the source host system; retrieving, by the failover host system, metadata pertaining to the memory region from the source host system via a remote direct memory access (RDMA) connection, wherein the metadata identifies one or more portions of the memory region, and wherein the RDMA connection was previously established between a network interface controller (NIC) of the failover host system and a NIC of the source host system before the failure; and for each of the one or more portions identified by the metadata, copying, by the failover host system, content of the portion from the memory region to the new memory region via the RDMA connection.
2 . The method of claim 1 further comprising, prior to the allocating:
receiving a virtual machine (VM) migrated from the source host system in response to the failure.
3 . The method of claim 2 wherein the allocating comprises:
detecting that a virtual persistent memory module exists in a configuration file of the VM, the virtual persistent memory module being mapped to the memory region; and
allocating the new memory region to have a same size as the virtual persistent memory module.
4 . The method of claim 3 further comprising, after the copying:
mapping the virtual persistent memory module to the new memory region.
5 . The method of claim 1 wherein the failure causes an operating system (OS) or hypervisor of the source host system to become inoperable.
6 . The method of claim 1 further comprising, prior to retrieving the metadata:
retrieving a baseline copy of the memory region from a storage backend shared by the source host system and the failover host system, the baseline copy representing a copy of the memory region as captured via a periodic flushing operation to the storage backend prior to the failure; and
populating the new memory region with the baseline copy.
7 . The method of claim 6 wherein the one or more portions of the memory region identified by the metadata include data updates absent in the baseline copy.
8 . A non-transitory computer readable storage medium having stored thereon program code executable by a failover host system, the program code embodying a method comprising, upon occurrence of a failure at a source host system:
allocating a new memory region in a physical memory of the failover host system that corresponds to a memory region in a physical memory of the source host system; retrieving metadata pertaining to the memory region from the source host system via a remote direct memory access (RDMA) connection, wherein the metadata identifies one or more portions of the memory region, and wherein the RDMA connection was previously established between a network interface controller (NIC) of the failover host system and a NIC of the source host system before the failure; and for each of the one or more portions identified by the metadata, copying content of the portion from the memory region to the new memory region via the RDMA connection.
9 . The non-transitory computer readable storage medium of claim 8 wherein the method further comprises, prior to the allocating:
receiving a virtual machine (VM) migrated from the source host system in response to the failure.
10 . The non-transitory computer readable storage medium of claim 9 wherein the allocating comprises:
detecting that a virtual persistent memory module exists in a configuration file of the VM, the virtual persistent memory module being mapped to the memory region; and
allocating the new memory region to have a same size as the virtual persistent memory module.
11 . The non-transitory computer readable storage medium of claim 10 wherein the method further comprises, after the copying:
mapping the virtual persistent memory module to the new memory region.
12 . The non-transitory computer readable storage medium of claim 8 wherein the failure causes an operating system (OS) or hypervisor of the source host system to become inoperable.
13 . The non-transitory computer readable storage medium of claim 8 wherein the method further comprises, prior to retrieving the metadata:
retrieving a baseline copy of the memory region from a storage backend shared by the source host system and the failover host system, the baseline copy representing a copy of the memory region as captured via a periodic flushing operation to the storage backend prior to the failure; and
populating the new memory region with the baseline copy.
14 . The non-transitory computer readable storage medium of claim 13 wherein the one or more portions of the memory region identified by the metadata include data updates absent in the baseline copy.
15 . A host system comprising:
a processor; a physical memory; a network interface controller (NIC); and a non-transitory computer readable medium having stored thereon program code that, when executed by the processor, causes the processor to, upon occurrence of a failure at another host system:
allocate a new memory region in the physical memory that corresponds to a memory region in a physical memory of said another host system;
retrieve metadata pertaining to the memory region from said another host system via a remote direct memory access (RDMA) connection, wherein the metadata identifies one or more portions of the memory region, and wherein the RDMA connection was previously established between the NIC of the host system and a NIC of said another host system before the failure; and
for each of the one or more portions identified by the metadata, copying content of the portion from the memory region to the new memory region via the RDMA connection.
16 . The host system of claim 15 wherein the program code further causes the processor to, prior to the allocating:
receive a virtual machine (VM) migrated from said another host system in response to the failure.
17 . The host system of claim 16 wherein the program code that causes the processor to allocate the new memory region comprises program code that causes the processor to:
detect that a virtual persistent memory module exists in a configuration file of the VM, the virtual persistent memory module being mapped to the memory region; and
allocate the new memory region to have a same size as the virtual persistent memory module.
18 . The host system of claim 17 wherein the program code further causes the processor to, after the copying:
map the virtual persistent memory module to the new memory region.
19 . The host system of claim 15 wherein the failure causes an operating system (OS) or hypervisor of said another host system to become inoperable.
20 . The host system of claim 15 wherein the program code further causes the processor to, prior to retrieving the metadata:
retrieve a baseline copy of the memory region from a storage backend shared by the host system and said another host system, the baseline copy representing a copy of the memory region as captured via a periodic flushing operation to the storage backend prior to the failure; and
populate the new memory region with the baseline copy.
21 . The host system of claim 20 wherein the one or more portions of the memory region identified by the metadata include data updates absent in the baseline copy.Join the waitlist — get patent alerts
Track US2023315593A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.