296
In contrast, the performance of memory-bound workloads is generally limited by
the memory subsystem and the long latencies of fetching data from caches and system
memory. An example is an application that randomly accesses data from data structures
in DRAM. In this case, adding more compute resources would not improve such an
application. Adding persistent memory to improve performance is usually an option for
memory-bound workloads as opposed to compute-bound workloads. Memory-bound
workloads usually have lower CPU utilization than compute-bound workloads, exhibit
CPU stalls due to memory transfers, and have high memory bandwidth.
Memory Latency vs. Memory Capacity
This concept is essential when discussing persistent memory. For this discussion, we
assume that DRAM access latencies are lower than persistent memory and that the
persistent memory capacity within the system is larger than DRAM. Workloads bound by
memory capacity can benefit from adding persistent memory in a volatile mode, while
workloads that are bound by memory latency are less likely to benefit.
Read vs. Write Performance
While each persistent memory technology is unique, it is important to understand
that there is usually a difference in the performance of reads (loads) vs. writes (stores).
Different media types exhibit varying degrees of asymmetric read-write performance
characteristics, where reads are generally much faster than writes. Therefore,
understanding the mix of loads and stores in an application workload is important for
understanding and optimizing performance.
Memory Access Patterns
A memory access pattern is the pattern with which a system or application reads and
writes to or from the memory. Memory hardware usually relies on temporal locality
(accessing recently used data) and spatial locality (accessing contiguous memory
addresses) for best performance. This is often achieved through some structure of fast
internal caches and intelligent prefetchers. The access pattern and level of locality can
drastically affect cache performance and can also have implications on parallelism
and distributions of workloads within shared memory systems. Cache coherency can
Chapter 15 profiling and performanCe
Précédent

- 320/457

Suivant