363
directly into the final persistent memory location, whereas remote replication to an SSD
requires an RDMA Write into the DRAM on the remote server, followed by a second local
DMA operation to move the remote write data from volatile DRAM into the final storage
location on the SSD or other legacy block storage device.
The performance challenge with replicating data to remote persistent memory is that
while large block sizes of 512KiB or larger can achieve good performance, as the size of
the writes being replicated gets smaller, the network overhead becomes a larger portion
of the total latency, and performance can suffer.
If the persistent memory is being used as an SSD replacement, the typical native
block storage size is 4K, avoiding some of the inefficiencies seen with small transfers.
If the persistent memory replaces a traditional SSD and data is written remotely to the
SSD, the latency improvements with persistent memory can be 10x or more.
The synchronous replication model implemented in librpmem means that small
data structures and pointer updates in local persistent memory result in small, very
inefficient, RDMA Writes followed by a small RDMA Read or Send to make that small
amount of write data persistent. This results in significant performance degradation
compared to writing only to local persistent memory. It makes the replication
performance very dependent on the local persistent memory write sequences, which
is heavily dependent on the application workload. In general, the larger the average
request size and the lower the number of rpmem_persist() calls that are required for a
given workload will improve the overall latency required for guaranteeing that data is
persistent.
It is possible to follow multiple RDMA Writes with single RDMA Read or Send
to make all of the preceding writes persistent. This reduces the impact of the size of
RDMA Writes on the overall performance of the proposed solution. But using this
mitigation, remember you are not guaranteed that any of the RDMA Writes is persistent
until RDMA Read completion returns or you receive RDMA Send with a confirmation.
The implementation that allows this approach is implemented in rpmem_flush() and
rpmem_drain() API call pair, where rpmem_flush() performs RDMA Write and returns
immediately and rpmem_drain() posts RDMA Read and waits for its completion (at the
time of publication it is not implemented in the write/send model).
There are many performance considerations, including the high-level networking
model being used. Traditional best-in-class networking architecture typically relies
on a pull model between the initiator and target. In a pull model, the initiator requests
resources from the target, but the target server only pulls the data across via RDMA
Chapter 18 remote persistent memory
directly into the final persistent memory location, whereas remote replication to an SSD
requires an RDMA Write into the DRAM on the remote server, followed by a second local
DMA operation to move the remote write data from volatile DRAM into the final storage
location on the SSD or other legacy block storage device.
The performance challenge with replicating data to remote persistent memory is that
while large block sizes of 512KiB or larger can achieve good performance, as the size of
the writes being replicated gets smaller, the network overhead becomes a larger portion
of the total latency, and performance can suffer.
If the persistent memory is being used as an SSD replacement, the typical native
block storage size is 4K, avoiding some of the inefficiencies seen with small transfers.
If the persistent memory replaces a traditional SSD and data is written remotely to the
SSD, the latency improvements with persistent memory can be 10x or more.
The synchronous replication model implemented in librpmem means that small
data structures and pointer updates in local persistent memory result in small, very
inefficient, RDMA Writes followed by a small RDMA Read or Send to make that small
amount of write data persistent. This results in significant performance degradation
compared to writing only to local persistent memory. It makes the replication
performance very dependent on the local persistent memory write sequences, which
is heavily dependent on the application workload. In general, the larger the average
request size and the lower the number of rpmem_persist() calls that are required for a
given workload will improve the overall latency required for guaranteeing that data is
persistent.
It is possible to follow multiple RDMA Writes with single RDMA Read or Send
to make all of the preceding writes persistent. This reduces the impact of the size of
RDMA Writes on the overall performance of the proposed solution. But using this
mitigation, remember you are not guaranteed that any of the RDMA Writes is persistent
until RDMA Read completion returns or you receive RDMA Send with a confirmation.
The implementation that allows this approach is implemented in rpmem_flush() and
rpmem_drain() API call pair, where rpmem_flush() performs RDMA Write and returns
immediately and rpmem_drain() posts RDMA Read and waits for its completion (at the
time of publication it is not implemented in the write/send model).
There are many performance considerations, including the high-level networking
model being used. Traditional best-in-class networking architecture typically relies
on a pull model between the initiator and target. In a pull model, the initiator requests
resources from the target, but the target server only pulls the data across via RDMA
Chapter 18 remote persistent memory
