Storage system with consistent termination of data replication across multiple distributed processing modules
Abstract
A first storage system in one illustrative embodiment is configured to participate in a replication process with a second storage system. A first processing module of a distributed storage controller of the first storage system detects a replication failure condition for a given write request received from a host device, and provides a corresponding notification to a second processing module of the distributed storage controller. The second processing module, responsive to receipt of the notification, instructs the first processing module and a plurality of additional processing modules of a same type as the first processing module to suspend generation of replication acknowledgments for write requests received from the host device. Responsive to receipt of confirmation from the first and additional processing modules of their suspended generation of replication acknowledgements, the second processing module instructs the first and additional processing modules to terminate the replication process.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. An apparatus comprising:
a first storage system comprising a plurality of storage nodes;
the first storage system being configured to participate in a replication process with a second storage system;
each of the storage nodes of the first storage system comprising a plurality of storage devices;
each of the storage nodes of the first storage system further comprising a set of processing modules configured to communicate over one or more networks with corresponding sets of processing modules on other ones of the storage nodes;
the sets of processing modules of the storage nodes collectively comprising at least a portion of a distributed storage controller of the first storage system;
wherein in conjunction with the replication process, a first one of the processing modules is configured to detect a replication failure condition for a given write request received from a host device, and to provide a notification to a second one of the processing modules of the detected replication failure condition;
the second processing module being configured, responsive to receipt of the notification of the detected replication failure condition, to instruct the first processing module and a plurality of additional ones of the processing modules of a same type as the first processing module to suspend generation of replication acknowledgments for write requests received from the host device;
the second processing module being further configured, responsive to receipt of confirmation from the first and additional processing modules of their suspended generation of replication acknowledgements, to instruct the first and additional processing modules to terminate the replication process;
wherein each of the storage nodes is implemented using at least one processing device comprising a processor coupled to a memory.
2. The apparatus of claim 1 wherein the first and second storage systems comprise respective content addressable storage systems having respective sets of non-volatile memory storage devices.
3. The apparatus of claim 1 wherein the first and second storage systems are associated with respective source and target sites of the replication process and wherein the source site comprises a production site data center and the target site comprises a disaster recovery site data center.
4. The apparatus of claim 1 wherein the replication process comprises a synchronous replication process in which write requests directed by the host device to the first storage system are mirrored to the second storage system.
5. The apparatus of claim 1 wherein the replication failure condition for the given write request comprises a failure to receive in the first processing module a response from the second storage system indicating that the given write request has been successfully mirrored to the second storage system.
6. The apparatus of claim 1 wherein the first and additional processing modules are each configured to set a replication barrier responsive to receipt of the instruction from the second processing module to suspend generation of replication acknowledgments for write requests received from the host device.
7. The apparatus of claim 1 wherein:
each of the sets of processing modules comprises one or more control modules;
at least one of the sets of processing modules comprises a management module;
the first processing module comprises a given one of the control modules;
the second processing module comprises the management module; and
the additional processing modules of the same type as the first processing module comprise respective additional ones of the control modules.
8. The apparatus of claim 7 wherein the management module comprises a system-wide management module of the first storage system.
9. The apparatus of claim 7 wherein the management module instructs all of the control modules of the distributed storage controller to terminate the replication process only after confirmation of suspended generation of replication acknowledgements is received from each of those control modules.
10. The apparatus of claim 7 wherein the management module instructs all of the control modules of the distributed storage controller to terminate the replication process by instructing each of the control modules to stop mirroring write requests to the second storage system.
11. The apparatus of claim 10 wherein different ones of the control modules stop mirroring write requests to the second storage system at different times.
12. The apparatus of claim 6 wherein each of the first and additional processing modules releases its set replication barrier in conjunction with termination of the replication process.
13. The apparatus of claim 1 wherein the first processing module acknowledges the given write request to the host device in conjunction with termination of the replication process.
14. The apparatus of claim 1 wherein the notification of the detected replication failure condition is one of a plurality of such notifications received in the second processing module from respective ones of the first and additional processing modules.
15. A method comprising:
configuring a first storage system to include a plurality of storage nodes each having a plurality of storage devices, each of the storage nodes further comprising a set of processing modules configured to communicate over one or more networks with corresponding sets of processing modules on other ones of the storage nodes;
configuring the first storage system to participate in a replication process with a second storage system; and
in conjunction with the replication process, a first one of the processing modules detecting a replication failure condition for a given write request received from a host device, and providing a notification to a second one of the processing modules of the detected replication failure condition;
the second processing module, responsive to receipt of the notification of the detected replication failure condition, instructing the first processing module and a plurality of additional ones of the processing modules of a same type as the first processing module to suspend generation of replication acknowledgments for write requests received from the host device;
the second processing module, responsive to receipt of confirmation from the first and additional processing modules of their suspended generation of replication acknowledgements, instructing the first and additional processing modules to terminate the replication process;
wherein the method is implemented by at least one processing device comprising a processor coupled to a memory.
16. The method of claim 15 wherein the replication failure condition for the given write request comprises a failure to receive in the first processing module a response from the second storage system indicating that the given write request has been successfully mirrored to the second storage system.
17. The method of claim 15 wherein:
each of the sets of processing modules comprises one or more control modules;
at least one of the sets of processing modules comprises a management module;
the first processing module comprises a given one of the control modules;
the second processing module comprises the management module; and
the additional processing modules of the same type as the first processing module comprise respective additional ones of the control modules.
18. A computer program product comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes said at least one processing device:
to configure a first storage system to include a plurality of storage nodes each having a plurality of storage devices, each of the storage nodes further comprising a set of processing modules configured to communicate over one or more networks with corresponding sets of processing modules on other ones of the storage nodes;
to configure the first storage system to participate in a replication process with a second storage system; and
in conjunction with the replication process, a first one of the processing modules being configured to detect a replication failure condition for a given write request received from a host device, and to provide a notification to a second one of the processing modules of the detected replication failure condition;
the second processing module being configured, responsive to receipt of the notification of the detected replication failure condition, to instruct the first processing module and a plurality of additional ones of the processing modules of a same type as the first processing module to suspend generation of replication acknowledgments for write requests received from the host device;
the second processing module being further configured, responsive to receipt of confirmation from the first and additional processing modules of their suspended generation of replication acknowledgements, to instruct the first and additional processing modules to terminate the replication process.
19. The computer program product of claim 18 wherein the replication failure condition for the given write request comprises a failure to receive in the first processing module a response from the second storage system indicating that the given write request has been successfully mirrored to the second storage system.
20. The computer program product of claim 18 wherein:
each of the sets of processing modules comprises one or more control modules;
at least one of the sets of processing modules comprises a management module;
the first processing module comprises a given one of the control modules;
the second processing module comprises the management module; and
the additional processing modules of the same type as the first processing module comprise respective additional ones of the control modules.Join the waitlist — get patent alerts
Track US10338851B1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.