314
Everything is built on top of libpmem and its persistence primitives that the library
uses to transfer data to persistent memory and persist it. Those primitives are also
exposed through libpmemobj-specific APIs to applications that wish to perform low- level
operations on persistent memory, such as manual cache flushing. These APIs are exposed
so the high-level library can instrument, intercept, and augment all stores to persistent
memory. This is useful for the instrumentation of runtime analysis tools such as Valgrind
pmemcheck, described in Chapter 12. More importantly, these functions are interception
points for data replication, both local and remote.
Replication is implemented in a way that ensures all data written prior to calling
drain will be safely stored in the replica as configured. A drain operation is a barrier that
waits for hardware buffers to complete their flush operation to ensure all writes have
reached the media. This works by initiating a write to the replica when a memory copy
or a flush is performed and then waits for those writes to finish in the drain call. This
mechanism guarantees the same behavior and ordering semantics for replicated and
non-replicated pools.
On top of persistence primitives provided by libpmem is an abstraction for fail-safe
modification of transactional data called unified logging. The unified log is a single
data structure and API for two different logging types used throughout libpmemobj to
ensure fail-safety: transactions and atomic operations. This is one of the most crucial,
performance-sensitive modules in the library because it is the hot code path of almost
every API. The unified log is a hybrid DRAM and persistent memory data structure
accessed through a runtime context that organizes all memory operations that need
to be performed within a single fail-safe atomic transaction and allows for critical
performance optimizations.
The persistent memory allocator operates in the unified log context of either a
transaction or a single atomic operation. This is the largest and most complex module in
libpmemobj and is used to manage the potentially large amounts of persistent memory
associated with the memory pool.
Each object stored in a persistent memory pool is represented by an object handle
of type PMEMoid (persistent memory object identifier). In practice, such a handle is
a unique object identifier (OID) of global scope, which means that two objects from
different pools will never have the same OID. An OID cannot be used as a direct pointer
to an object. Each time the program attempts to read or write object data, it must obtain
the current memory address of the object by converting its OID into a pointer. In contrast
to the memory address, the OID value for a given object does not change during the
life of an object, except for a realloc(), and remains valid after closing and reopening
Chapter 16 pMDK Internals: IMportant algorIthMs anD Data struCtures
Everything is built on top of libpmem and its persistence primitives that the library
uses to transfer data to persistent memory and persist it. Those primitives are also
exposed through libpmemobj-specific APIs to applications that wish to perform low- level
operations on persistent memory, such as manual cache flushing. These APIs are exposed
so the high-level library can instrument, intercept, and augment all stores to persistent
memory. This is useful for the instrumentation of runtime analysis tools such as Valgrind
pmemcheck, described in Chapter 12. More importantly, these functions are interception
points for data replication, both local and remote.
Replication is implemented in a way that ensures all data written prior to calling
drain will be safely stored in the replica as configured. A drain operation is a barrier that
waits for hardware buffers to complete their flush operation to ensure all writes have
reached the media. This works by initiating a write to the replica when a memory copy
or a flush is performed and then waits for those writes to finish in the drain call. This
mechanism guarantees the same behavior and ordering semantics for replicated and
non-replicated pools.
On top of persistence primitives provided by libpmem is an abstraction for fail-safe
modification of transactional data called unified logging. The unified log is a single
data structure and API for two different logging types used throughout libpmemobj to
ensure fail-safety: transactions and atomic operations. This is one of the most crucial,
performance-sensitive modules in the library because it is the hot code path of almost
every API. The unified log is a hybrid DRAM and persistent memory data structure
accessed through a runtime context that organizes all memory operations that need
to be performed within a single fail-safe atomic transaction and allows for critical
performance optimizations.
The persistent memory allocator operates in the unified log context of either a
transaction or a single atomic operation. This is the largest and most complex module in
libpmemobj and is used to manage the potentially large amounts of persistent memory
associated with the memory pool.
Each object stored in a persistent memory pool is represented by an object handle
of type PMEMoid (persistent memory object identifier). In practice, such a handle is
a unique object identifier (OID) of global scope, which means that two objects from
different pools will never have the same OID. An OID cannot be used as a direct pointer
to an object. Each time the program attempts to read or write object data, it must obtain
the current memory address of the object by converting its OID into a pointer. In contrast
to the memory address, the OID value for a given object does not change during the
life of an object, except for a realloc(), and remains valid after closing and reopening
Chapter 16 pMDK Internals: IMportant algorIthMs anD Data struCtures
