374
On a NUMA system, the greater the distance between a processor and a memory
bank, the slower the processor’s access to that memory bank. Performance-sensitive
applications should therefore be configured so they allocate memory from the closest
possible memory bank.
Performance-sensitive applications should also be configured to execute on a set
number of cores, particularly in the case of multithreaded applications. Because firstlevel caches are usually small, if multiple threads execute on one core, each thread
will potentially evict cached data accessed by a previous thread. When the operating
system attempts to multitask between these threads, and the threads continue to evict
each other’s cached data, a large percentage of their execution time is spent on cache
line replacement. This issue is referred to as cache thrashing. We therefore recommend
that you bind a multithreaded application to a NUMA node rather than a single core,
since this allows the threads to share cache lines on multiple levels (first-, second-, and
last-level cache) and minimizes the need for cache fill operations. However, binding
an application to a single core may be performant if all threads are accessing the
same cached data. numactl allows you to bind an application to a particular core or
NUMA node and to allocate the memory associated with a core or set of cores to that
application.
NUMACTL Linux Utility
On Linux we can use the numactl utility to display the NUMA hardware configuration
and control which cores and threads application processes can run. The libnuma library
included in the numactl package offers a simple programming interface to the NUMA
policy supported by the kernel. It is useful for more fine-grained tuning than the numactl
utility. Further information is available in the numa(7) man page.
Figure 19-1. A two-socket CPU NUMA architecture showing local and remote
memory access
Chapter 19 advanCed topiCs
Précédent

- 396/457

Suivant