138
significant threat to both primary and uninterruptible power supply (Li
et al. 2013). Human causes include human error or malicious attacks from
insiders or external third parties. The latter is largely a security issue while
the former is a training and behavioural one. Li et al. (2013) document a
wide range of public cloud outages resulting from human error including
vehicle accidents, power shutdowns, and inputting commands in error. As
discussed earlier in this section, application and system level failures can be
technological causes of full service outage. In these instances, for application failures, the key requirement is business continuity through redundancy and rollback. It should be noted that a number of middleware
approaches have been applied to address application-level reliability via
application-independent failure detection, checkpoint and rollback and
recovery (e.g. Hormati, et al. 2014), optimal replica placement (e.g. An
et al. 2014), stop and copy VM migration (Sampaio and Barbosa 2018),
and entity reputation management (Abawajy 2011). For system level failures, the primary focus is minimising recovery time (Singh et al. 2016). It
is important to note that while these causes are isolated, they may be cascading, natural causes can result in unanticipated technological failures,
which in turn may be exacerbated by human errors, and so forth.
As discussed in Chap. 2, the SLA details the level of service to be provided, often in the form of specific QoS metrics (Ghazizadeh and Cusack
2018). Obviously, in the context of trust, there is a close relationship
between SLA metrics and monitoring, and unsurprisingly this is a major
focus of both cloud monitoring systems (see Fatema et al. 2014) and
trustworthy cloud computing research. This research primarily focuses on
the decomposition of SLA parameters in to low-level system performance
metrics, mapping these in to KPIs, and then ultimately aggregating these
KPIs in to some form of aggregated quality indicator that can be used to
mitigate transactional risk (Sun et al. 2012). A wide range of techniques
are used to measure and predict cloud service performance (and indeed
SLA violation). Typical metrics include availability, bandwidth, cost
(including energy), CPU cycle, service duration, memory, request arrival
rate, space/storage. Upgrade request frequency as well as other more specific performance metrics (throughput, response time, execution time
etc.) are also present, although the importance of these will vary by cloud
service (Faniyi and Bahsoon 2015). Cloud service providers may also
include metrics that specifically acknowledge the risk of failure e.g. the
maximum fraction of SLA violations allowed or penalty rates (Faniyi and
Bahsoon 2015). Notably, security is an attribute metric that is extremely
O. M. ALOFE AND K. FATEMA
significant threat to both primary and uninterruptible power supply (Li
et al. 2013). Human causes include human error or malicious attacks from
insiders or external third parties. The latter is largely a security issue while
the former is a training and behavioural one. Li et al. (2013) document a
wide range of public cloud outages resulting from human error including
vehicle accidents, power shutdowns, and inputting commands in error. As
discussed earlier in this section, application and system level failures can be
technological causes of full service outage. In these instances, for application failures, the key requirement is business continuity through redundancy and rollback. It should be noted that a number of middleware
approaches have been applied to address application-level reliability via
application-independent failure detection, checkpoint and rollback and
recovery (e.g. Hormati, et al. 2014), optimal replica placement (e.g. An
et al. 2014), stop and copy VM migration (Sampaio and Barbosa 2018),
and entity reputation management (Abawajy 2011). For system level failures, the primary focus is minimising recovery time (Singh et al. 2016). It
is important to note that while these causes are isolated, they may be cascading, natural causes can result in unanticipated technological failures,
which in turn may be exacerbated by human errors, and so forth.
As discussed in Chap. 2, the SLA details the level of service to be provided, often in the form of specific QoS metrics (Ghazizadeh and Cusack
2018). Obviously, in the context of trust, there is a close relationship
between SLA metrics and monitoring, and unsurprisingly this is a major
focus of both cloud monitoring systems (see Fatema et al. 2014) and
trustworthy cloud computing research. This research primarily focuses on
the decomposition of SLA parameters in to low-level system performance
metrics, mapping these in to KPIs, and then ultimately aggregating these
KPIs in to some form of aggregated quality indicator that can be used to
mitigate transactional risk (Sun et al. 2012). A wide range of techniques
are used to measure and predict cloud service performance (and indeed
SLA violation). Typical metrics include availability, bandwidth, cost
(including energy), CPU cycle, service duration, memory, request arrival
rate, space/storage. Upgrade request frequency as well as other more specific performance metrics (throughput, response time, execution time
etc.) are also present, although the importance of these will vary by cloud
service (Faniyi and Bahsoon 2015). Cloud service providers may also
include metrics that specifically acknowledge the risk of failure e.g. the
maximum fraction of SLA violations allowed or penalty rates (Faniyi and
Bahsoon 2015). Notably, security is an attribute metric that is extremely
O. M. ALOFE AND K. FATEMA
