this ApplicationMaster allocates Containers from NodeManager for MapReduce
task. This approach reduces the load on ResourceManager and distributes it across
ApplicationMasters on the datanodes for each job. This way, using YARN the
hadoop cluster can grow up to 10,000 nodes. Earlier benchmark without YARN on
Hadoop 1.x was up to 4000 nodes. This way YARN provides scalability to
the Hadoop cluster along with different programming platforms to be incorporated
in hadoop framework.
5 Big data Tools Development for Drug Discovery
There have been efforts by various scientific groups to use HPC, grid technologies
for drug discovery. Multiple docking tools like DOCK6 [28], Gold [29], Autodock
Vina [30], and some others are already available in the parallel mode on HPC
platform. Most of these tools are fast and robust; however, they have their own
scoring functions based on molecular mechanics force fields and other geometrical
descriptors. Although, improvements are still going on in enhancing the scoring
function and guiding it further toward higher efficiency and accuracy. Docking with
the concept of flexible ligand and protein still remains to be time-consuming calculation. Docking of multiple ligands to single protein or multiple ligands with
multiple proteins may be some of the future challenges in docking area.
Understanding the flexibility of both the proteins and ligands has been taken care by
some of the currently available molecular simulation packages like AMBER [31],
CHARMM [32], GROMACS [33], and NAMD [34]. All these packages are known
to be scalable on the HPC platform. Although molecular simulations are
time-consuming, they still prove to be the best in understanding the allowed flexibility of proteins, ligands, active sites, and other biomolecular entities. The advent
of cloud and big data technologies promises to accelerate the drug development
process using MapReduce [27] and spark methods coupled with machine learning
and deep learning analytics. The tools like DIVE [35], HiMach [36], and HTMD
[37] have been developed for molecular simulations as well as trajectory visualization and analysis. Many more tools may be getting developed using these newer
technologies.
Bioinformatics group at C-DAC, Pune, has been addressing the issue on data
analytics and visualization of trajectories in structural biology domain using HPC
technologies combined with big data technologies. Various analytics tools have been
developed and tested on Hadoop platform using MapReduce as shown in Fig. 7. At
this stage, analytics tools for multiple molecular trajectories include hydrogen bond
calculations, identifying water molecules and bridged water-mediated interactions.
Other big data analytics tools for RMSD, 2DRMSD, RMSF, water density,
WHAM-based free energy calculations are in the process of development. Few of the
big data analytics tools which have been already developed proved to be useful in the
process of drug discovery. These tools have described below.
356
R. R. Joshi et al.
Précédent

- 364/413

Suivant