scalable, multi-user platform for big data analytics in the field of Bioinformatics.
The use of big data in the Bioinformatics is an emerging field which presents new
opportunities to medical researchers and paves the way toward prediction of personalized medicines. The greatest challenge lies in designing a strategy to acquire
the data followed by filtering it to meet the appropriate decision-making demands.
This can be achieved by bringing together experts from clinical medicines,
computer science, bioinformatics, biotechnology, and statistics and address the
challenge of the data management and analytics solutions toward precision biology.
Hadoop [26]-based platform with MapReduce and spark-based algorithms may be
useful to make all the analysis optimized with fast calculation. Hadoop- and
MapReduce [27]-based algorithms implemented on scalable architecture have been
discussed further along with drug repurposing big data case study for cancer protein.
4 Big data Technology Components
Hadoop
Apache Hadoop is an open-source software framework for storage and large-scale
processing of datasets on clusters of commodity hardware. Hadoop has gained lots of
popularity among the peer parallel data processing tools because of its simplicity,
efficiency, cost, and reliability. Hadoop can be built on the commodity hardware.
Hadoop has major three components. Hadoop Distributed File System (HDFS),
YARN scheduler and resource negotiating framework and the MapReduce [27]
programming framework. A typical framework of Hadoop test bed is shown in Fig. 3.
Fig. 3 Basic architecture diagram of hadoop test bed
352
R. R. Joshi et al.
Précédent

- 360/413

Suivant