343
ACPI NFIT Health Event Notification
Due to the potential loss of quality of service, operating systems and privileged
applications may not want to actively poll persistent memory devices to retrieve device
health. Thus, the ACPI specification has defined a passive notification method to allow
the persistent memory device to notify when a significant change in device health
has occurred. Persistent memory device vendors and platform BIOS vendors decide
which device health changes are significant enough to trigger an NVDIMM Firmware
Interface Table (NFIT) health event notification. Upon receipt of an NFIT health event,
a notification to the operating system is expected to call an _NCH or a _DSM attached to
the persistent memory device and take appropriate action based on the data returned.
Unsafe/Dirty Shutdown
An unsafe or dirty shutdown on persistent memory means that the persistent memory
device power-down sequence or platform power-down sequence may have failed
to write all in-flight data from the system’s persistence domain to persistent media.
(Chapter 2 describes persistence domains.) A dirty shutdown is expected to be a very
rare event, but they can happen due to a variety of reasons such as physical hardware
issues, power spikes, thermal events, and so on.
A persistent memory device does not know if any application data was lost as a result
of the incomplete power-down sequence. It can only detect if a series of events occurred
in which data may have been lost. In the best-case scenario, there might not have been
any applications that were in the process of writing data when the dirty shutdown
occurred.
The RAS mechanism described here requires the platform BIOS and persistent
memory vendor to maintain a persistent rolling counter that is incremented anytime a
dirty shutdown is detected. The ACPI specification refers to such a mechanism as the Data
Loss Count (DLC) that can be returned as part of the Get NVDIMM Boot Status(_NBS)
persistent memory device method.
Referring to the output from ndctl in Listing 17-1, the "shutdown_count" is reported
in the health information. Similarly, the output from ipmctl in Listing 17-2 reports
"LatchedDirtyShutdownCount" as the dirty shutdown counter. For both outputs, a value
of 1 means no issues were detected.
Chapter 17 reliability, availability, and ServiCeability (raS)
ACPI NFIT Health Event Notification
Due to the potential loss of quality of service, operating systems and privileged
applications may not want to actively poll persistent memory devices to retrieve device
health. Thus, the ACPI specification has defined a passive notification method to allow
the persistent memory device to notify when a significant change in device health
has occurred. Persistent memory device vendors and platform BIOS vendors decide
which device health changes are significant enough to trigger an NVDIMM Firmware
Interface Table (NFIT) health event notification. Upon receipt of an NFIT health event,
a notification to the operating system is expected to call an _NCH or a _DSM attached to
the persistent memory device and take appropriate action based on the data returned.
Unsafe/Dirty Shutdown
An unsafe or dirty shutdown on persistent memory means that the persistent memory
device power-down sequence or platform power-down sequence may have failed
to write all in-flight data from the system’s persistence domain to persistent media.
(Chapter 2 describes persistence domains.) A dirty shutdown is expected to be a very
rare event, but they can happen due to a variety of reasons such as physical hardware
issues, power spikes, thermal events, and so on.
A persistent memory device does not know if any application data was lost as a result
of the incomplete power-down sequence. It can only detect if a series of events occurred
in which data may have been lost. In the best-case scenario, there might not have been
any applications that were in the process of writing data when the dirty shutdown
occurred.
The RAS mechanism described here requires the platform BIOS and persistent
memory vendor to maintain a persistent rolling counter that is incremented anytime a
dirty shutdown is detected. The ACPI specification refers to such a mechanism as the Data
Loss Count (DLC) that can be returned as part of the Get NVDIMM Boot Status(_NBS)
persistent memory device method.
Referring to the output from ndctl in Listing 17-1, the "shutdown_count" is reported
in the health information. Similarly, the output from ipmctl in Listing 17-2 reports
"LatchedDirtyShutdownCount" as the dirty shutdown counter. For both outputs, a value
of 1 means no issues were detected.
Chapter 17 reliability, availability, and ServiCeability (raS)
