Various tools discussed earlier in this chapter have been used for parallel
visualization and efficient and fast analysis of Ras docking and simulation trajectories data. In-house computational facility BRAF has been used where these tools
are already deployed and tested. The results would help the experimentalist to select
the better ligand for further steps of drug development.
7 Latest Development in Big data
Bioinformatics is a technology-driven science. There have been major technological shifts which are driving the data-driven science. With the ever-increasing data,
the storage and analysis of huge data are becoming very tedious and most of the
data remains unanalyzed. For example, the sequencing of genomes of various
organisms is generating petabytes to zetabytes of data. Also, the development of
new sequencing technology like nanopore is capable of producing long reads
generating huge data [71]. The assembly of such genomes put out a huge challenge
on the Big Data technologies. The Apache Hadoop has also enhanced to tackle such
challenges like Yarn which allows different data processing engines including graph
processing, stream processing as well as batch processing. The MapReduce
framework provided by Apache Hadoop is good for batch processing. In case of
iterative processing where the data need to be read many times, the MapReduce is
not efficient. MapReduce relies heavily on disk input/output so it is slow. The
Apache Spark addresses this limitation of Hadoop and provides in memory computing but reducing disk input/output. Spark supports in memory computing and
optimizes disk performance by lazy loading and cache mechanism. Hence, spark is
suitable for iterative computing.
Recent progressions have empowered the most precision analytics strategies at
the “single cell” level. The sequencing of single cell brings about enormous volume
and complexities of information and presents an extraordinary chance to comprehend the cell level heterogeneity. The latest developments highlight the inherent
opportunities and challenges in Big Data analytics. The recently created technologies like erasure encoding mechanism [72] in Hadoop 3.x tend to resolve the
difficulties postured by several big data problems like single cell transcriptome
analysis in bioinformatics and present great opportunity to develop cutting-edge
technologies for the future research problems. The HDFS uses redundancy for high
availability of data. It provides great benefit at the cost of storage byte. Generally,
with replication factor of 3, HDFS uses three times more storage data redundancy.
So it is very costly in terms of storage. The erasure encoding mechanism in Hadoop
3.x provides same storage safety at the cost of 50% storage overhead. This is
effective when data is more and its access frequency is less.
370
R. R. Joshi et al.
visualization and efficient and fast analysis of Ras docking and simulation trajectories data. In-house computational facility BRAF has been used where these tools
are already deployed and tested. The results would help the experimentalist to select
the better ligand for further steps of drug development.
7 Latest Development in Big data
Bioinformatics is a technology-driven science. There have been major technological shifts which are driving the data-driven science. With the ever-increasing data,
the storage and analysis of huge data are becoming very tedious and most of the
data remains unanalyzed. For example, the sequencing of genomes of various
organisms is generating petabytes to zetabytes of data. Also, the development of
new sequencing technology like nanopore is capable of producing long reads
generating huge data [71]. The assembly of such genomes put out a huge challenge
on the Big Data technologies. The Apache Hadoop has also enhanced to tackle such
challenges like Yarn which allows different data processing engines including graph
processing, stream processing as well as batch processing. The MapReduce
framework provided by Apache Hadoop is good for batch processing. In case of
iterative processing where the data need to be read many times, the MapReduce is
not efficient. MapReduce relies heavily on disk input/output so it is slow. The
Apache Spark addresses this limitation of Hadoop and provides in memory computing but reducing disk input/output. Spark supports in memory computing and
optimizes disk performance by lazy loading and cache mechanism. Hence, spark is
suitable for iterative computing.
Recent progressions have empowered the most precision analytics strategies at
the “single cell” level. The sequencing of single cell brings about enormous volume
and complexities of information and presents an extraordinary chance to comprehend the cell level heterogeneity. The latest developments highlight the inherent
opportunities and challenges in Big Data analytics. The recently created technologies like erasure encoding mechanism [72] in Hadoop 3.x tend to resolve the
difficulties postured by several big data problems like single cell transcriptome
analysis in bioinformatics and present great opportunity to develop cutting-edge
technologies for the future research problems. The HDFS uses redundancy for high
availability of data. It provides great benefit at the cost of storage byte. Generally,
with replication factor of 3, HDFS uses three times more storage data redundancy.
So it is very costly in terms of storage. The erasure encoding mechanism in Hadoop
3.x provides same storage safety at the cost of 50% storage overhead. This is
effective when data is more and its access frequency is less.
370
R. R. Joshi et al.
