All these programs have modules written for performing different analyses like
RMSD, RMSF, radius of gyration, PCA [40], distance calculations, H-bond analysis, and MMGBSA [41] free energy calculations. However, many of these programs are either inefficient or very slow in calculating the H-bond interactions
within solute and especially between the solute and the solvent (water molecules).
These programs are highly time-consuming and also have constraint in dealing with
the large size data for example 500 GB or beyond. This drawback of the existing
tools suggests a strong need for the development of water-mediated H-bond analysis tool which is capable of handling a very large size of trajectories and also be
executed parallel to reduce the time. The water molecules added to the system may
play a crucial role in the activity or functioning of that particular molecule. Hence,
understanding the role and mechanism of such water molecules and their interactions with the solute (protein/RNA/DNA or drug) molecules is very important [42,
43]. In order to achieve this, a big data analytics tool for hydrogen bond calculation
was developed by Bioinformatics group C-DAC.
The MapReduce algorithm for H-bond calculation was developed and ported/
tested on Hadoop cluster. The algorithm flow has been shown in Fig. 8a for H-bond
calculation using the MapReduce approach. The HDFS file system was used to
store the multiple molecular trajectories data. The current version of tool can
analyze trajectory data in the PDB format generated using molecular dynamics
packages like AMBER [31], GROMACS [33], CHARMM [32]. The tool is scalable or portable on any distributed computing platform and can find out H-bonds
between all types of residues including water. However, the tool requires a significant amount of time for executing the preprocessing stage where, the PDB files
are generated from the trajectories and copied on the distributed HDFS storage.
Despite this overhead, the overall performance of the tool is better than currently
existing tools such as CPPTRAJ or PTRAJ [38], especially for trajectories with a
large number of water molecules. The benchmarking of H-bond tool is shown in
Fig. 8b. The benchmarking of up to 5.5 TB data is carried out, and it shows near
linear scale up. Additionally, the tool can also help identify water-mediated interactions such as water bridges easily.
5.2 Molecular Conformation Generation on Cloud
(MOSAIC)
Drug databases usually contain millions of ligands, and for each ligand, there can be
billions of conformations [44, 45]. Such billions of conformations need to be
docked on to a target which is a generally a protein molecule. Generation and
optimization of such billions of ligand conformations is a huge computational
problem, since it involves the use of advanced methods like molecular mechanics,
semi-empirical and quantum techniques [46, 47]. The application of an embarrassingly parallel approach accompanied by virtualized resource scaling and an
358
R. R. Joshi et al.
Précédent

- 366/413

Suivant