297
also affect multiprocessor performance, which means that certain memory access
patterns place a ceiling on parallelism. Many well-defined memory access patterns exist,
including but not limited to sequential, strided, linear, and random.
It is much easier to measure, control, and optimize memory accesses on systems that
run only one application. In the cloud and virtualized environments, applications within
the guests can be running any type of application and workload, including web servers,
databases, or an application server. This makes it much harder to ensure memory
accesses are fully optimized for the hardware as the access patterns are essentially
random.
I/O Storage Bound Workloads
A program is I/O bound if it would go faster if the I/O subsystem were faster. We are
primarily interested in the block-based disk I/O subsystem here, but it could also include
other subsystems such as the network. An I/O bound state is undesirable because
it means that the CPU must stall its operation while waiting for data to be loaded or
unloaded from main memory or storage. Depending on where the data is and the
latency of the storage device, this can invoke a voluntary context switching of the current
application thread with another. A voluntary context switch occurs when a thread blocks
because it requires a resource that is not immediately available or takes a long time
to respond. With faster computation speed being the primary goal of each successive
computer generation, there is a strong imperative to avoid I/O bound states. Eliminating
them can often yield a more economic improvement in performance than upgrading the
CPU or memory.
Determining the Suitability of Workloads
for Persistent Memory
Persistent memory technologies may not solve every workload performance problem.
You should understand the workload and platform on which it is currently running
when considering persistent memory. As a simple example, consider a computeintensive workload that relies heavily on floating-point arithmetic. The performance of
this application is likely limited by the floating-point unit in the CPU and not any part
of the memory subsystem. In that case, adding persistent memory to the platform will
likely have little impact on this application’s performance. Now consider an application
Chapter 15 profiling and performanCe
also affect multiprocessor performance, which means that certain memory access
patterns place a ceiling on parallelism. Many well-defined memory access patterns exist,
including but not limited to sequential, strided, linear, and random.
It is much easier to measure, control, and optimize memory accesses on systems that
run only one application. In the cloud and virtualized environments, applications within
the guests can be running any type of application and workload, including web servers,
databases, or an application server. This makes it much harder to ensure memory
accesses are fully optimized for the hardware as the access patterns are essentially
random.
I/O Storage Bound Workloads
A program is I/O bound if it would go faster if the I/O subsystem were faster. We are
primarily interested in the block-based disk I/O subsystem here, but it could also include
other subsystems such as the network. An I/O bound state is undesirable because
it means that the CPU must stall its operation while waiting for data to be loaded or
unloaded from main memory or storage. Depending on where the data is and the
latency of the storage device, this can invoke a voluntary context switching of the current
application thread with another. A voluntary context switch occurs when a thread blocks
because it requires a resource that is not immediately available or takes a long time
to respond. With faster computation speed being the primary goal of each successive
computer generation, there is a strong imperative to avoid I/O bound states. Eliminating
them can often yield a more economic improvement in performance than upgrading the
CPU or memory.
Determining the Suitability of Workloads
for Persistent Memory
Persistent memory technologies may not solve every workload performance problem.
You should understand the workload and platform on which it is currently running
when considering persistent memory. As a simple example, consider a computeintensive workload that relies heavily on floating-point arithmetic. The performance of
this application is likely limited by the floating-point unit in the CPU and not any part
of the memory subsystem. In that case, adding persistent memory to the platform will
likely have little impact on this application’s performance. Now consider an application
Chapter 15 profiling and performanCe
