310
references an address, the system translates the virtual address to a physical address.
The physical address points to memory physically connected to a CPU. Chapter 19
describes exactly how this operation works and shows why high-capacity memory
systems can benefit from using large or huge pages provided by the operating system.
A common practice in software is to have most of the data allocations done when the
application starts. Operating systems try to allocate memory associated with the CPU on
which the thread executes. The operating system scheduler then tries to always schedule
the thread on a CPU that it last ran in the hopes that the data still remains in one of the
CPU caches. On a multi-socket system, this may result in all the objects being allocated
in the memory of a single socket, which can create NUMA performance issues. Accessing
data on a remote CPU incurs a latency performance penalty.
Some applications delay reserving memory until the data is accessed for the first
time. This can alleviate some NUMA issues. It is important to understand how your
workload allocates data to understand the NUMA performance.
Thread Migration
Thread migration, which is the movement of software threads across sockets by the
operating system scheduler, is the most common cause of NUMA issues. Once objects
are allocated in memory, accessing them from another physical CPU from which they
were originally allocated incurs a latency penalty. Even though you may allocate your
data on a socket where the accessing thread is currently running, unless you have
specific affinity bindings or other safeguards, the thread may move to any other core or
socket in the future. You can track thread migration by identifying which cores threads
are running on and which sockets those cores belong to. Figure 15-10 shows an example
of this analysis from VTune Profiler.
Figure 15-10. VTune Profiler identifying thread migration across cores and
sockets (packages)
Chapter 15 profiling and performanCe
references an address, the system translates the virtual address to a physical address.
The physical address points to memory physically connected to a CPU. Chapter 19
describes exactly how this operation works and shows why high-capacity memory
systems can benefit from using large or huge pages provided by the operating system.
A common practice in software is to have most of the data allocations done when the
application starts. Operating systems try to allocate memory associated with the CPU on
which the thread executes. The operating system scheduler then tries to always schedule
the thread on a CPU that it last ran in the hopes that the data still remains in one of the
CPU caches. On a multi-socket system, this may result in all the objects being allocated
in the memory of a single socket, which can create NUMA performance issues. Accessing
data on a remote CPU incurs a latency performance penalty.
Some applications delay reserving memory until the data is accessed for the first
time. This can alleviate some NUMA issues. It is important to understand how your
workload allocates data to understand the NUMA performance.
Thread Migration
Thread migration, which is the movement of software threads across sockets by the
operating system scheduler, is the most common cause of NUMA issues. Once objects
are allocated in memory, accessing them from another physical CPU from which they
were originally allocated incurs a latency penalty. Even though you may allocate your
data on a socket where the accessing thread is currently running, unless you have
specific affinity bindings or other safeguards, the thread may move to any other core or
socket in the future. You can track thread migration by identifying which cores threads
are running on and which sockets those cores belong to. Figure 15-10 shows an example
of this analysis from VTune Profiler.
Figure 15-10. VTune Profiler identifying thread migration across cores and
sockets (packages)
Chapter 15 profiling and performanCe
