6 Big Data
333
data. When a patient is admitted to ICU, devices reading the patient’s vitals send a
new data record every second. Each ICU patient generates a large amount of data,
which is often times very noisy. This data is traditionally summarized into 15 minute
or longer intervals due to either the systems’ inability to store and process the large
amounts of more granular data or the cost-effectiveness of doing so. This approach
limits the types of signal detection and analysis that can be performed on the data
at scale. For example, determining which medicines bring down fever fastest would
require a fine-grained measure of various streams (heart rate, breathing, movement,
pain, etc.). Determining the efficacy of the medicine and enabling decision making
in real time or near real time would require data time resolution of seconds or
minutes. By implementing the data lake, researchers at Mercy Hospital were able to
capture diagnostic data at high temporal rates, which in turn enabled real-time and
near-real-time processing of the data.
This real-time data-on-demand model for researchers and clinicians is enabled by
a combination of Sqoop, Storm, and HBase for more granular updates. In addition,
Hive has provided a SQL-like approach to enabling the scalability of the Hadoop
data lake.
6.4 An Introduction to Spark: An Innovative Paradigm
in Big Data
One fundamental component to the Big Data Ecosystem not yet mentioned is Spark
[25]. Although Hadoop captures the most attention for distributed data analytics,
there are alternatives that provide advantages to the typical Hadoop platform.
Apache Spark is an open source cluster computing framework originally developed
in the AMPLab at the University of California, Berkeley, but was later donated to
the Apache Software Foundation where it remains today [26].
Spark is a scalable data analytics platform that incorporates primitives for inmemory computing and typically demonstrates a significant speedup of the classic
Hadoop’s cluster compute and storage approach. Spark is implemented in the Scala
programming language and provides a unique environment for large-scale data
storage and processing. In contrast to Hadoop’s two-stage, disk-based MapReduce
paradigm, Spark’s multi-stage, in-memory primitives approach provides performance up to 100 times faster for certain applications [25]. By allowing user
programs to load data into a cluster’s memory and query it repeatedly, Spark can
enable a fast, efficient implementation of a variety of machine learning algorithms.
Similar to traditional Hadoop system, Spark requires a cluster manager and a
distributed storage system. For cluster management, Spark supports the standalone
native Spark cluster, Hadoop YARN or Apache Mesos. For distributed storage,
Spark can interface with a wide variety of systems including HDFS, Cassandra,
OpenStack Swift, Amazon S3, or custom solutions.
333
data. When a patient is admitted to ICU, devices reading the patient’s vitals send a
new data record every second. Each ICU patient generates a large amount of data,
which is often times very noisy. This data is traditionally summarized into 15 minute
or longer intervals due to either the systems’ inability to store and process the large
amounts of more granular data or the cost-effectiveness of doing so. This approach
limits the types of signal detection and analysis that can be performed on the data
at scale. For example, determining which medicines bring down fever fastest would
require a fine-grained measure of various streams (heart rate, breathing, movement,
pain, etc.). Determining the efficacy of the medicine and enabling decision making
in real time or near real time would require data time resolution of seconds or
minutes. By implementing the data lake, researchers at Mercy Hospital were able to
capture diagnostic data at high temporal rates, which in turn enabled real-time and
near-real-time processing of the data.
This real-time data-on-demand model for researchers and clinicians is enabled by
a combination of Sqoop, Storm, and HBase for more granular updates. In addition,
Hive has provided a SQL-like approach to enabling the scalability of the Hadoop
data lake.
6.4 An Introduction to Spark: An Innovative Paradigm
in Big Data
One fundamental component to the Big Data Ecosystem not yet mentioned is Spark
[25]. Although Hadoop captures the most attention for distributed data analytics,
there are alternatives that provide advantages to the typical Hadoop platform.
Apache Spark is an open source cluster computing framework originally developed
in the AMPLab at the University of California, Berkeley, but was later donated to
the Apache Software Foundation where it remains today [26].
Spark is a scalable data analytics platform that incorporates primitives for inmemory computing and typically demonstrates a significant speedup of the classic
Hadoop’s cluster compute and storage approach. Spark is implemented in the Scala
programming language and provides a unique environment for large-scale data
storage and processing. In contrast to Hadoop’s two-stage, disk-based MapReduce
paradigm, Spark’s multi-stage, in-memory primitives approach provides performance up to 100 times faster for certain applications [25]. By allowing user
programs to load data into a cluster’s memory and query it repeatedly, Spark can
enable a fast, efficient implementation of a variety of machine learning algorithms.
Similar to traditional Hadoop system, Spark requires a cluster manager and a
distributed storage system. For cluster management, Spark supports the standalone
native Spark cluster, Hadoop YARN or Apache Mesos. For distributed storage,
Spark can interface with a wide variety of systems including HDFS, Cassandra,
OpenStack Swift, Amazon S3, or custom solutions.
