4 Architecting IoT Cloud
217
Fig. 4.30 Data that is split
up into multiple blocks and
stored across a set of
connected nodes
File (e.g.,
sensor data)
…
…
Node 1
Node 2
Node n
4.6.3.3 Distributed File Systems
Data lake systems are mainly comprised of scalable data stores and distributed
file systems that deploy data by partitioning larger files into smaller blocks or by
partitioning tables that are appropriated across a pool of data servers (data nodes).
In addition, data is usually replicated in several nodes to increase fault tolerance.
In such a distributed system, a load balancer is utilized to increase scalability. This
technique is demonstrated in Fig. 4.30.
Apache Hadoop’s distributed file system and Amazon S3’s Cloud storage are
two well-known data lakes:
• Hadoop Distributed File System (HDFS): HDFS is a Java-based file system
created to be deployed across several commodity servers easily and provides
dependable and scalable data storage. The details of HDFS will be addressed in
Chap. 6.
• Amazon S3 Storage Service (Amazon S3): Amazon S3 Storage Service stores
data objects using an uncomplicated web service interface. It can store and fetch
data in any amount. It is utilized as the main storage option for Cloud-native
applications, as a bulk data depository or as a data lake for analytics. Amazon S3
is also useful for serverless computing, recovery, and backup.
4.6.3.4 Data Lake Tiers
Generally, it is highly recommended to engineer a data lake based on multiple
tiers/layers (not less than two), with one being a quarantine zone. This is particularly
important within tightly regulated industries that utilize highly sensitive data that
must be manually verified by data stewards before transfer to another zone with
wider user access. It is also highly suggested dividing the second tier into several
217
Fig. 4.30 Data that is split
up into multiple blocks and
stored across a set of
connected nodes
File (e.g.,
sensor data)
…
…
Node 1
Node 2
Node n
4.6.3.3 Distributed File Systems
Data lake systems are mainly comprised of scalable data stores and distributed
file systems that deploy data by partitioning larger files into smaller blocks or by
partitioning tables that are appropriated across a pool of data servers (data nodes).
In addition, data is usually replicated in several nodes to increase fault tolerance.
In such a distributed system, a load balancer is utilized to increase scalability. This
technique is demonstrated in Fig. 4.30.
Apache Hadoop’s distributed file system and Amazon S3’s Cloud storage are
two well-known data lakes:
• Hadoop Distributed File System (HDFS): HDFS is a Java-based file system
created to be deployed across several commodity servers easily and provides
dependable and scalable data storage. The details of HDFS will be addressed in
Chap. 6.
• Amazon S3 Storage Service (Amazon S3): Amazon S3 Storage Service stores
data objects using an uncomplicated web service interface. It can store and fetch
data in any amount. It is utilized as the main storage option for Cloud-native
applications, as a bulk data depository or as a data lake for analytics. Amazon S3
is also useful for serverless computing, recovery, and backup.
4.6.3.4 Data Lake Tiers
Generally, it is highly recommended to engineer a data lake based on multiple
tiers/layers (not less than two), with one being a quarantine zone. This is particularly
important within tightly regulated industries that utilize highly sensitive data that
must be manually verified by data stewards before transfer to another zone with
wider user access. It is also highly suggested dividing the second tier into several
