24
Intel Machine Instructions for Persistent Memory
Applicable to Intel- and AMD-based ADR platforms, executing an Intel 64 and 32 architecture
store instruction is not enough to make data persistent since the data may be sitting in the
CPU caches indefinitely and could be lost by a power failure. Additional cache flush actions
are required to make the stores persistent. Importantly, these non- privileged cache flush
operations can be called from user space, meaning applications decide when and where to
fence and flush data. Table 2-1 summarizes each of these instructions. For more detailed
information, the Intel 64 and 32 Architectures Software Developer Manuals are online at
https://software.intel.com/en-us/articles/intel-sdm.
Developers should primarily focus on CLWB and Non-Temporal Stores if available
and fall back to the others as necessary. Table 2-1 lists other opcodes for completeness.
Table 2-1. Intel architecture instructions for persistent memory
OPCODE
Description
CLFLUSH
this instruction, supported in many generations of Cpu, flushes a single
cache line. historically, this instruction is serialized, causing multiple CLFLush
instructions to execute one after the other, without any concurrency.
CLFLUSHOPT
(followed by an
SFENCE)
this instruction, newly introduced for persistent memory support, is like
CLFLUSH but without the serialization. to flush a range, the software executes a
CLFLUSHOPT instruction for each 64-byte cache line in the range, followed by a
single SFENCE instruction to ensure the flushes are complete before continuing.
CLFLUSHOPT is optimized, hence the name, to allow some concurrency when
executing multiple CLFLUSHOPT instructions back-to-back.
CLWB (followed by
an SFENCE)
the effect of cache line writeback (CLWB) is the same as CLFLUSHOPT except
that the cache line may remain valid in the cache but is no longer dirty since it
was flushed. this makes it more likely to get a cache hit on this line if the data
is accessed again later.
Non-temporal
stores (followed
by an SFENCE)
this feature has existed for a while in x86 Cpus. these stores are “write
combining” and bypass the Cpu cache; using them does not require a flush. a
final SFENCE instruction is still required to ensure the stores have reached the
persistence domain.
(continued)
Chapter 2 persistent MeMory arChiteCture
Intel Machine Instructions for Persistent Memory
Applicable to Intel- and AMD-based ADR platforms, executing an Intel 64 and 32 architecture
store instruction is not enough to make data persistent since the data may be sitting in the
CPU caches indefinitely and could be lost by a power failure. Additional cache flush actions
are required to make the stores persistent. Importantly, these non- privileged cache flush
operations can be called from user space, meaning applications decide when and where to
fence and flush data. Table 2-1 summarizes each of these instructions. For more detailed
information, the Intel 64 and 32 Architectures Software Developer Manuals are online at
https://software.intel.com/en-us/articles/intel-sdm.
Developers should primarily focus on CLWB and Non-Temporal Stores if available
and fall back to the others as necessary. Table 2-1 lists other opcodes for completeness.
Table 2-1. Intel architecture instructions for persistent memory
OPCODE
Description
CLFLUSH
this instruction, supported in many generations of Cpu, flushes a single
cache line. historically, this instruction is serialized, causing multiple CLFLush
instructions to execute one after the other, without any concurrency.
CLFLUSHOPT
(followed by an
SFENCE)
this instruction, newly introduced for persistent memory support, is like
CLFLUSH but without the serialization. to flush a range, the software executes a
CLFLUSHOPT instruction for each 64-byte cache line in the range, followed by a
single SFENCE instruction to ensure the flushes are complete before continuing.
CLFLUSHOPT is optimized, hence the name, to allow some concurrency when
executing multiple CLFLUSHOPT instructions back-to-back.
CLWB (followed by
an SFENCE)
the effect of cache line writeback (CLWB) is the same as CLFLUSHOPT except
that the cache line may remain valid in the cache but is no longer dirty since it
was flushed. this makes it more likely to get a cache hit on this line if the data
is accessed again later.
Non-temporal
stores (followed
by an SFENCE)
this feature has existed for a while in x86 Cpus. these stores are “write
combining” and bypass the Cpu cache; using them does not require a flush. a
final SFENCE instruction is still required to ensure the stores have reached the
persistence domain.
(continued)
Chapter 2 persistent MeMory arChiteCture
