337
Unconsumed uncorrectable error handling may be implemented differently on
different vendor platforms, but at the core, there will always be a mechanism to discover
the unconsumed uncorrectable error, a mechanism to signal the operating system of an
unconsumed uncorrectable error, and a mechanism for the operating system to query
information about the unconsumed uncorrectable error. As shown in Figure 17-2, these
three mechanisms work together to proactively keep the operating system informed of
all discovered uncorrectable errors during runtime.
Patrol Scrub
Patrol scrub (also known as memory scrubbing) is a long-standing RAS feature for
volatile memory that can also be extended to persistent memory. It is an excellent
example of how a platform can discover uncorrectable errors in the background during
normal operation.
Patrol scrubbing is done using a hardware engine, on either the platform or on the
memory device, which generates requests to memory addresses on the memory device.
The engine generates memory requests at a predefined frequency. Given enough time,
it will eventually access every memory address. The frequency in which patrol scrub
generates requests produces no noticeable impact on the memory device’s quality of
service.
Figure 17-2. Unconsumed uncorrectable error handling
Chapter 17 reliability, availability, and ServiCeability (raS)
Unconsumed uncorrectable error handling may be implemented differently on
different vendor platforms, but at the core, there will always be a mechanism to discover
the unconsumed uncorrectable error, a mechanism to signal the operating system of an
unconsumed uncorrectable error, and a mechanism for the operating system to query
information about the unconsumed uncorrectable error. As shown in Figure 17-2, these
three mechanisms work together to proactively keep the operating system informed of
all discovered uncorrectable errors during runtime.
Patrol Scrub
Patrol scrub (also known as memory scrubbing) is a long-standing RAS feature for
volatile memory that can also be extended to persistent memory. It is an excellent
example of how a platform can discover uncorrectable errors in the background during
normal operation.
Patrol scrubbing is done using a hardware engine, on either the platform or on the
memory device, which generates requests to memory addresses on the memory device.
The engine generates memory requests at a predefined frequency. Given enough time,
it will eventually access every memory address. The frequency in which patrol scrub
generates requests produces no noticeable impact on the memory device’s quality of
service.
Figure 17-2. Unconsumed uncorrectable error handling
Chapter 17 reliability, availability, and ServiCeability (raS)
