385
We can also create a number of possible configurations if we were to create multiple
namespaces per region or partition the /dev/pmem* devices using fdisk or parted, for
example. Doing this provides greater flexibility and isolation of the resulting logical
volumes. However, if a physical NVDIMM fails, the impact is significantly greater since it
would impact some or all of the file systems depending on the configuration.
Creating complex RAID volume groups may protect the data but at the cost of not
efficiently using all the persistent memory capacity for data. Additionally, complex RAID
volume groups do not support the DAX feature that some applications may require.
The mmap( ) MAP_SYNC Flag
Introduced in the Linux kernel v4.15, the MAP_SYNC flag ensures that any needed file
system metadata writes are completed before a process is allowed to modify directly
mapped data. The MAP_SYNC flag was added to the mmap() system call to request the
synchronous behavior; in particular, the guarantee provided by this flag is
While a block is writeably mapped into page tables of this mapping, it is
guaranteed to be visible in the file at that offset also after a crash.
This means the file system will not silently relocate the block, and it will ensure that the
file’s metadata is in a consistent state so that the blocks in question will be present after
a crash. This is done by ensuring that any needed metadata writes were done before the
process is allowed to write pages affected by that metadata.
When a persistent memory region is mapped using MAP_SYNC, the memory
management code will check to see whether there are metadata writes pending for the
affected file. However, it will not actually flush those writes out. Instead, the pages are
mapped read only with a special flag, forcing a page fault when the process first attempts
to perform a write to one of those pages. The fault handler will then synchronously flush
out any dirty metadata, set the page permissions to allow the write, and return. At that
point, the process can write the page safely, since all the necessary metadata changes
have already made it to persistent storage.
The result is a relatively simple mechanism that will perform far better than
the currently available alternative of manually calling fsync() before each write to
persistent memory. The additional IO from fsync() can potentially cause the process to
block in what was supposed to be a simple memory write, introducing latency that may
be unexpected and unwanted.
Chapter 19 advanCed topiCs
We can also create a number of possible configurations if we were to create multiple
namespaces per region or partition the /dev/pmem* devices using fdisk or parted, for
example. Doing this provides greater flexibility and isolation of the resulting logical
volumes. However, if a physical NVDIMM fails, the impact is significantly greater since it
would impact some or all of the file systems depending on the configuration.
Creating complex RAID volume groups may protect the data but at the cost of not
efficiently using all the persistent memory capacity for data. Additionally, complex RAID
volume groups do not support the DAX feature that some applications may require.
The mmap( ) MAP_SYNC Flag
Introduced in the Linux kernel v4.15, the MAP_SYNC flag ensures that any needed file
system metadata writes are completed before a process is allowed to modify directly
mapped data. The MAP_SYNC flag was added to the mmap() system call to request the
synchronous behavior; in particular, the guarantee provided by this flag is
While a block is writeably mapped into page tables of this mapping, it is
guaranteed to be visible in the file at that offset also after a crash.
This means the file system will not silently relocate the block, and it will ensure that the
file’s metadata is in a consistent state so that the blocks in question will be present after
a crash. This is done by ensuring that any needed metadata writes were done before the
process is allowed to write pages affected by that metadata.
When a persistent memory region is mapped using MAP_SYNC, the memory
management code will check to see whether there are metadata writes pending for the
affected file. However, it will not actually flush those writes out. Instead, the pages are
mapped read only with a special flag, forcing a page fault when the process first attempts
to perform a write to one of those pages. The fault handler will then synchronously flush
out any dirty metadata, set the page permissions to allow the write, and return. At that
point, the process can write the page safely, since all the necessary metadata changes
have already made it to persistent storage.
The result is a relatively simple mechanism that will perform far better than
the currently available alternative of manually calling fsync() before each write to
persistent memory. The additional IO from fsync() can potentially cause the process to
block in what was supposed to be a simple memory write, introducing latency that may
be unexpected and unwanted.
Chapter 19 advanCed topiCs
