136
• The impact of system failure should be minimised for individual
users, the overall number of users affected, and the downtime associated for the failure;
• Service performance and capacity should be maximised to reduce the
impact of reduced performance even if no failure is detected; and,
• Business continuity should be maximised by responding to failures
when they occur, protecting the integrity of data, and recovering as
soon as possible.
Reliability and high availability are closely related and regarded as significant challenges in cloud computing. Obviously, cloud service providers
and scholars invest a significant amount of effort in to the design of faulttolerant, attack-resilient and reliable systems. A detailed discussion of this
is beyond the scope of this chapter. These innovations are often opaque to
the user. As such, we provide a high-level overview of approaches to reliability including ensuring reliability by design through monitoring, redundancy and disaster recovery, and the evaluation of performance and quality
of service (QoS).
A major focus of computer science research is reliability by design so
that no one point of failure can result in the failure of the entire system.
There are a wide variety of causes of unplanned cloud outages including
infrastructure or software failures, planning mistakes, human error, or
external attacks (Endo et al. 2017). Three main strategies are employed to
counter such failures namely, monitoring, redundancy, and disaster recovery. In the terminology of trust, two could be classified as trust-building
mechanisms (monitoring and redundancy) while the third, disaster recovery, could be classified as a trust repair mechanism. A wide variety of general purpose and vendor-specific monitoring tools are used in cloud
computing. From the user perspective, these are primarily used for
accounting and billing, security and privacy assurance, and SLA management, while for the cloud service provider they may be used for other
reliability functions, for example fault management (Fatema et al. 2014).
As mentioned earlier, gray failures may not be detectable by extant monitoring systems that focus on singular failure detection. To mitigate the risk
of such failures, Huang et al. (2017) suggest that cloud service providers
must move to multi-dimensional cloud health monitoring. While accepting monitoring all applications and workloads in hyperscale multi-tenant
systems is not feasible, they propose a number of techniques to close the
observation gap including approximating application views, aggregating
O. M. ALOFE AND K. FATEMA
• The impact of system failure should be minimised for individual
users, the overall number of users affected, and the downtime associated for the failure;
• Service performance and capacity should be maximised to reduce the
impact of reduced performance even if no failure is detected; and,
• Business continuity should be maximised by responding to failures
when they occur, protecting the integrity of data, and recovering as
soon as possible.
Reliability and high availability are closely related and regarded as significant challenges in cloud computing. Obviously, cloud service providers
and scholars invest a significant amount of effort in to the design of faulttolerant, attack-resilient and reliable systems. A detailed discussion of this
is beyond the scope of this chapter. These innovations are often opaque to
the user. As such, we provide a high-level overview of approaches to reliability including ensuring reliability by design through monitoring, redundancy and disaster recovery, and the evaluation of performance and quality
of service (QoS).
A major focus of computer science research is reliability by design so
that no one point of failure can result in the failure of the entire system.
There are a wide variety of causes of unplanned cloud outages including
infrastructure or software failures, planning mistakes, human error, or
external attacks (Endo et al. 2017). Three main strategies are employed to
counter such failures namely, monitoring, redundancy, and disaster recovery. In the terminology of trust, two could be classified as trust-building
mechanisms (monitoring and redundancy) while the third, disaster recovery, could be classified as a trust repair mechanism. A wide variety of general purpose and vendor-specific monitoring tools are used in cloud
computing. From the user perspective, these are primarily used for
accounting and billing, security and privacy assurance, and SLA management, while for the cloud service provider they may be used for other
reliability functions, for example fault management (Fatema et al. 2014).
As mentioned earlier, gray failures may not be detectable by extant monitoring systems that focus on singular failure detection. To mitigate the risk
of such failures, Huang et al. (2017) suggest that cloud service providers
must move to multi-dimensional cloud health monitoring. While accepting monitoring all applications and workloads in hyperscale multi-tenant
systems is not feasible, they propose a number of techniques to close the
observation gap including approximating application views, aggregating
O. M. ALOFE AND K. FATEMA
