309
These performance metrics help to determine if memory is the bottleneck in your
application, and if so, which level of the memory hierarchy is the most impactful. Many
tools can pinpoint source code locations and memory objects responsible for the
bottleneck. If persistent memory is the bottleneck, review the “Guided Data Placement”
section to ensure that persistent memory is being used efficiently. Performance
optimizations like cache blocking, software prefetching, and improved memory access
patterns may also help relieve bottlenecks in the memory hierarchy. You must determine
how to refactor the software to more efficiently use memory, and metrics like these can
point you in the right direction.
NUMA Optimizations
NUMA-related performance issues were described in the “Characterizing the Workload”
section; we discuss NUMA in more detail in Chapter 19. If you identify performance
issues related to NUMA memory accesses, two things should be considered: data
allocation vs. first access, and thread migration.
Data Allocation vs. First Access
Data allocation is the process of allocating or reserving some amount of virtual address
space for an object. The virtual address space for a process is the set of virtual memory
addresses that it can use. The address space for each process is private and cannot be
accessed by other processes unless it is shared. A virtual address does not represent
the actual physical location of an object in memory. Instead, the system maintains a
multilayered page table, which is an internal data structure used to translate virtual
addresses into their corresponding physical addresses. Each time an application thread
Figure 15-9. VTune Profiler memory analysis of a workload showing a
breakdown of CPU cache, DRAM, and persistent memory accesses
Chapter 15 profiling and performanCe
These performance metrics help to determine if memory is the bottleneck in your
application, and if so, which level of the memory hierarchy is the most impactful. Many
tools can pinpoint source code locations and memory objects responsible for the
bottleneck. If persistent memory is the bottleneck, review the “Guided Data Placement”
section to ensure that persistent memory is being used efficiently. Performance
optimizations like cache blocking, software prefetching, and improved memory access
patterns may also help relieve bottlenecks in the memory hierarchy. You must determine
how to refactor the software to more efficiently use memory, and metrics like these can
point you in the right direction.
NUMA Optimizations
NUMA-related performance issues were described in the “Characterizing the Workload”
section; we discuss NUMA in more detail in Chapter 19. If you identify performance
issues related to NUMA memory accesses, two things should be considered: data
allocation vs. first access, and thread migration.
Data Allocation vs. First Access
Data allocation is the process of allocating or reserving some amount of virtual address
space for an object. The virtual address space for a process is the set of virtual memory
addresses that it can use. The address space for each process is private and cannot be
accessed by other processes unless it is shared. A virtual address does not represent
the actual physical location of an object in memory. Instead, the system maintains a
multilayered page table, which is an internal data structure used to translate virtual
addresses into their corresponding physical addresses. Each time an application thread
Figure 15-9. VTune Profiler memory analysis of a workload showing a
breakdown of CPU cache, DRAM, and persistent memory accesses
Chapter 15 profiling and performanCe
