I. HDFS
Hadoop Distributed File System (HDFS) is built to provide high-throughput,
reliable, efficient, and fault-tolerant file system. It can provide streaming reads and
writes for large files. The basic architecture diagram of HDFS is shown in Fig. 4.
As shown in the figure, the HDFS has two main components, namenode, and
datanode. HDFS is mainly designed for low-cost hardware, and hence, it can be
built on cluster of commodity hardware. In HDFS, the file is divided into fixed size
blocks or chunks of 128 MB each except the last chunk. The fixed size 128 MB can
be configured with various needs. Namenode contains the metadata information of
all the files. It stores information regarding the block of file stored on datanodes,
while datanodes actually store the block of data. Each block is stored on three
datanodes of the cluster. This policy provides reliability at the cost of redundancy.
Generally, two copies of blocks are stored on two different datanodes of the same
rack of cluster, while the third copy is stored on the datanodes of the different rack
of the same cluster. These two racks are connected by a very high-speed network
switch. This policy ensures the reliability of the HDFS file system. In case, if any
two nodes fail, still the data can be accessed from the datanode having this third
copy of the data. Datanodes periodically updates their state to the namenode so that
namenode can be aware of the overall state of cluster. While scheduling
MapReduce [27] job, the hadoop framework ensures with most possibility that the
mapper task should run on the same datanode where the actual data is residing. This
avoids significant network overhead. This policy of hadoop improves the performance of the overall cluster.
Fig. 4 Basic architecture diagram of Hadoop Distributed File System (HDFS)
Turbo Analytics: Applications of Big Data and HPC in Drug …
353
Hadoop Distributed File System (HDFS) is built to provide high-throughput,
reliable, efficient, and fault-tolerant file system. It can provide streaming reads and
writes for large files. The basic architecture diagram of HDFS is shown in Fig. 4.
As shown in the figure, the HDFS has two main components, namenode, and
datanode. HDFS is mainly designed for low-cost hardware, and hence, it can be
built on cluster of commodity hardware. In HDFS, the file is divided into fixed size
blocks or chunks of 128 MB each except the last chunk. The fixed size 128 MB can
be configured with various needs. Namenode contains the metadata information of
all the files. It stores information regarding the block of file stored on datanodes,
while datanodes actually store the block of data. Each block is stored on three
datanodes of the cluster. This policy provides reliability at the cost of redundancy.
Generally, two copies of blocks are stored on two different datanodes of the same
rack of cluster, while the third copy is stored on the datanodes of the different rack
of the same cluster. These two racks are connected by a very high-speed network
switch. This policy ensures the reliability of the HDFS file system. In case, if any
two nodes fail, still the data can be accessed from the datanode having this third
copy of the data. Datanodes periodically updates their state to the namenode so that
namenode can be aware of the overall state of cluster. While scheduling
MapReduce [27] job, the hadoop framework ensures with most possibility that the
mapper task should run on the same datanode where the actual data is residing. This
avoids significant network overhead. This policy of hadoop improves the performance of the overall cluster.
Fig. 4 Basic architecture diagram of Hadoop Distributed File System (HDFS)
Turbo Analytics: Applications of Big Data and HPC in Drug …
353
