355
system issues optimized flush instructions to flush each cache line in the list to the
persistence domain. This is followed by an SFENCE to guarantee these writes complete
before new writes are handled. At this point, the previous writes that were flushed in the
RDMA Send list are now persistent.
Performance Implications of the General-Purpose Remote
Replication Method
The general-purpose remote replication method requires that RDMA of the initiator
software follows a number of RDMA Write(s) with an RDMA Send. After the target NIC
finishes flushing the requested regions, an RDMA Send from the target goes back to
the initiator to affirm that the initiator application can consider those writes persistent.
This additional send/receive/send/receive messaging has an effect on latency and
throughput to make the writes persistent and has 50% higher latency than the appliance
remote replication method. The extra messaging has an effect on overall bandwidth and
scalability of all the RDMA connections running on those NICs.
Also, if the size of the RDMA Write that needs to be made persistent is small, the
efficiency of the connection drops dramatically as the extra messaging overhead
becomes a significant component of the overall latency. Additionally, the target
node CPU and caches are consumed for that operation. The same data is essentially
transmitted twice: once from NIC (via PCIe) to the CPU L3 cache and then from the CPU
L3 cache to the memory controller (iMC).
Appliance Remote Replication Method
Users of persistent memory on an Intel platform can use non-allocating write flows by
enabling the feature on the specific PCI root complex where incoming writes from the
NIC will enter into the CPU’s internal fabric and out to the persistent memory. Using the
non-allocating write flow, the incoming RDMA Writes will bypass CPU caches and go
directly to the persistence domain. This means that writes do not need to be flushed to
the persistence domain by the target system CPU.
The I/O pipeline still needs to be flushed to the persistence domain. This is more
efficiently accomplished by issuing a small RDMA Read to any memory address on the
same RDMA connection as the RDMA Writes; the memory address does not need to
be one that was written or is persistent. The RDMA specification clearly states that an
RDMA Read will force the previous RDMA Writes to complete first. This ordering rule is
Chapter 18 remote persistent memory
Précédent

- 378/457

Suivant