333
© The Author(s) 2020
S. Scargall, Programming Persistent Memory, https://doi.org/10.1007/978-1-4842-4932-1_17
CHAPTER 17
Reliability, Availability,
and Serviceability (RAS)
This chapter describes the high-level architecture of reliability, availability, and
serviceability (RAS) features designed for persistent memory. Persistent memory RAS
features were designed to support the unique error-handling strategy required for an
application when persistent memory is used. Error handling is an important part of the
program’s overall reliability, which directly affects the availability of applications. The
error-handling strategy for applications impacts what percentage of the expected time
the application is available to do its job.
Persistent memory vendors and platform vendors will both decide which RAS
features and how they will be implemented at the lowest hardware levels. Some
common RAS features were designed and documented in the ACPI specification, which
is maintained and owned by the UEFI Forum (https://uefi.org/). In this chapter,
we try to attain a general perspective of these ACPI-defined RAS features and call out
vendor-specific details if warranted.
Dealing with Uncorrectable Errors
The main memory of a server is protected using error correcting codes (ECC). This is
a common hardware feature that can automatically correct many memory errors that
happen due to transient hardware issues, such as power spikes, soft media errors, and so
on. If an error is severe enough, it will corrupt enough bits that ECC cannot correct; the
result is called an uncorrectable error (UE).
Uncorrectable errors in persistent memory require special RAS handling that differs
from how a platform may traditionally handle volatile memory uncorrectable errors.
© The Author(s) 2020
S. Scargall, Programming Persistent Memory, https://doi.org/10.1007/978-1-4842-4932-1_17
CHAPTER 17
Reliability, Availability,
and Serviceability (RAS)
This chapter describes the high-level architecture of reliability, availability, and
serviceability (RAS) features designed for persistent memory. Persistent memory RAS
features were designed to support the unique error-handling strategy required for an
application when persistent memory is used. Error handling is an important part of the
program’s overall reliability, which directly affects the availability of applications. The
error-handling strategy for applications impacts what percentage of the expected time
the application is available to do its job.
Persistent memory vendors and platform vendors will both decide which RAS
features and how they will be implemented at the lowest hardware levels. Some
common RAS features were designed and documented in the ACPI specification, which
is maintained and owned by the UEFI Forum (https://uefi.org/). In this chapter,
we try to attain a general perspective of these ACPI-defined RAS features and call out
vendor-specific details if warranted.
Dealing with Uncorrectable Errors
The main memory of a server is protected using error correcting codes (ECC). This is
a common hardware feature that can automatically correct many memory errors that
happen due to transient hardware issues, such as power spikes, soft media errors, and so
on. If an error is severe enough, it will corrupt enough bits that ECC cannot correct; the
result is called an uncorrectable error (UE).
Uncorrectable errors in persistent memory require special RAS handling that differs
from how a platform may traditionally handle volatile memory uncorrectable errors.
