352
from the initiator node to the persistent memory on the target node, that write or send
data needs to be flushed to the persistence domain on the remote system. Alternatively,
the remote write or send data needs to bypass CPU caches on the remote node to avoid
having to be flushed.
Different vendor-specific platform features add an extra challenge to RDMA and to
remote persistent memory. Intel platforms typically use a feature called allocating writes
or Direct Data IO (DDIO) which allows incoming writes to be placed directly into the
CPU’s L3 cache. The data is immediately visible to any application wanting to read the
data. However, having allocating writes enabled means that RDMA Writes to persistent
memory now have to be flushed to the persistence domain on the target node.
On Intel platforms, allocating writes can be disabled by turning on non-allocating
write I/O flows which forces the write data to bypass cache and be placed directly into
the persistent memory, governed by the location of the RDMA Write sink buffer. This
would slow down applications that will immediately touch the newly written data
because they incur the penalty to pull the data into CPU cache. However, this simplifies
making remote writes to persistent memory simpler and faster because cache flushing
on the remote target node can be avoided. An additional complication to using nonallocating write mode on an Intel platform is that an entire PCI root complex must be
enabled for this write mode. This means that any inbound writes that come through
that PCI root complex, for any device connected downstream of it, will have write-data
bypass CPU caches, causing possible additional performance latency as a side effect.
Intel specifies two methods for forcing writes to remote persistent memory into the
persistence domain:
1. A general-purpose remote replication method that does not rely
on Intel non- allocating write mode and assumes some or all of the
remote write data will end up in CPU cache on the target system
2. A high-performance appliance remote replication method that
uses the Intel platform-specific non-allocating write mode and
is probably more suited to an appliance product where there is
complete control over the hardware configuration to control what
is connected to which PCI root complex
Chapter 18 remote persistent memory
Précédent

- 375/457

Suivant