334
Persistent memory uncorrectable errors are persistent. Unlike volatile memory, if
power is lost or an application crashes and restarts, the uncorrectable error will remain
on the hardware. This can lead to an application getting stuck in an infinite loop such as
1. Application starts
2. Reads a memory address
3. Encounters uncorrectable error
4. Crashes (or system crashes and reboots)
5. Starts and resumes operation from where it left off
6. Performs a read on the same memory address that triggered the
previous restart
7. Crashes (or system crashes and reboots)
8. …
9. Repeats infinitely until manual intervention
The operating system and applications may need to address uncorrectable errors in
three main ways:
• When consuming previously undetected uncorrectable errors during
runtime
• When unconsumed uncorrectable errors are detected at runtime
• When mitigating uncorrectable memory locations detected at boot
Consumed Uncorrectable Error Handling
When an uncorrectable error is detected on a requested memory address, data
poisoning is used to inform the CPU that the data requested has an uncorrectable error.
When the hardware detects an uncorrectable memory error, it routes a poison bit along
with the data to the CPU. For the Intel architecture, when the CPU detects this poison
bit, it sends a processor interrupt signal to the operating system to notify it of this error.
This signal is called a machine check exception (MCE). The operating system can then
Chapter 17 reliability, availability, and ServiCeability (raS)
Précédent

- 357/457

Suivant