137
observations from multiple components to infer the likelihood of a gray
failure in an isolated component, as well as temporal analysis (Huang et al.
2017). As noted briefly in Chap. 1, monitoring data can be used more
widely in the context of building knowledge-based trust. Emeakaroha
et  al. (2016) have proposed a system and show through experimental
studies with business decision-makers that such monitoring systems can be
used to build trust through communication strategies such as trust labels
(Emeakaroha et al. 2016; van der Werff et al. 2018).
Cloud failures can be caused by issues that occur at different levels in
the cloud stack e.g. at the data, application, and/or system level (Huang
et al. 2017). Given organisational and consumer concerns about data and
availability of data in the event of a failure, it is unsurprising that in addition to general system redundancy, data redundancy is a primary concern
of cloud service providers. Data replication and erasure coding are commonly used data redundancy techniques in cloud computing (Nachiappan
et al. 2017). With simple data replication, data is replicated in at least two
locations on distributed cloud storage systems so that in the event of storage failure, it is just served from the replicated copy (Plank 2013). As such,
data loss only occurs if data corrupted on all storage targets the replicated
copies (Rajaasekharan 2014). As simple data replication carries a significant resource overhead in terms of storage, network and associated energy
consumption, hyperscale cloud service providers, such as Facebook and
Microsoft, use more advanced erasure coding, such as K out of N codes,
to detect and correct errors in cloud storage, and provide a less resource
intensive means to reconstruct data from parity data (Nachiappan et  al.
2017; Rajaasekharan 2014).
Disasters differ in terms of scale and impact (although this is subjective), and are typically unpredicted events that occur relatively rarely over
the lifetime of a given system. A full cloud service outage occurs more
frequently than one might imagine but due to the disaster recovery systems in place, the recovery time is extremely fast. Disasters can result from
natural, human, or technological causes, or a combination of two or more
of these (Singh et al. 2016). To mitigate the impact of natural disasters or
large-scale malicious physical attacks, cloud service providers, like many IT
organisations, use distributed backups, online and offline, in geographic
locations that are located sufficiently distant to avoid a homogenous natural event (Pokharel et  al. 2010). Maintaining two infrastructures is
extremely costly. However, cloud outages can also result from relatively
small-scale localised natural causes, for example lightning strikes are a
7 TRUSTWORTHY CLOUD COMPUTING
Précédent

- 155/166

Suivant