332
N. Balac
ZooKeeper and application-specific conventions. ZooKeeper recipes [22] provide
a simple service that can be used to build powerful abstractions.
6.3.8 Data Lakes and Warehouses
A data lake is defined as a centralized storage repository or system that contains a
massive amount of structured, semi-structured, and unstructured raw data. The data
structure and requirements are not defined until the data is needed to run different
types of analytics. Often times a data lake is a single store of all raw enterprise data
including copies of source system data, log files, clickstreams, social media, and
output from IoT devices. It can also include transformed data used for delivering
dashboards, reporting, visualization, new types of real-time analytics, and machine
learning.
A data warehouse is a database optimized to analyze relational data coming
from transactional systems and business applications. The data structure and schema
are defined in advance to optimize for fast queries. The data from a warehouse is
typically used for reporting and analysis. Data is cleaned, enriched, and transformed,
so it can act as the “single source of truth” that users can trust [23]. Many
organizations support both a data warehouse and a data lake, as they serve different
needs and use cases. A data lake enables storage of non-relational data from mobile
apps, IoT devices, and social media and does not require the schema to be defined
when data is captured (referred to as “schema on read”). Massive amounts of data
can be stored without careful schema design, thereby enabling flexibility in terms
of the kinds of questions or data analytics that might need to be performed in the
future. Data lakes enable different types of analytics including big data analytics,
text mining, real-time stream data processing, and machine learning.
One great, early example of a successful data lake is the Big Data implementation
at Mercy Hospital. The hospital leveraged technology to improve medical outcomes
for patients by utilizing one of the first comprehensive, integrated electronic health
record (EHR) systems to provide real-time, paperless access to patient information.
Utilizing the EHR from Epic Systems, every patient’s activity, including clinical
and financial interactions, was captured. The hospital needed to address several
typical challenges associated with Big Data implementations including scalability,
data schema requirements, and response to large data queries. Mercy, in partnership
with Hortonworks (one of the early Big Data providers at the time), has created
the Mercy Data Library, a Hadoop-based data lake. This data lake enabled the
integration, ingestion, and processing of large amounts of batch data extracts from
relational systems, real-time data directly from Epic HER, and information from
social media and even weather sources [24]. The combination of all of these data
sets in a common platform enables the hospital to ask and answer questions at that
were previously impossible due to scale, cost, or both.
One of the projects implemented on the Mercy system utilized thee advanced
analytics techniques on a large amount of intensive care unit (ICU) patients’ vitals
N. Balac
ZooKeeper and application-specific conventions. ZooKeeper recipes [22] provide
a simple service that can be used to build powerful abstractions.
6.3.8 Data Lakes and Warehouses
A data lake is defined as a centralized storage repository or system that contains a
massive amount of structured, semi-structured, and unstructured raw data. The data
structure and requirements are not defined until the data is needed to run different
types of analytics. Often times a data lake is a single store of all raw enterprise data
including copies of source system data, log files, clickstreams, social media, and
output from IoT devices. It can also include transformed data used for delivering
dashboards, reporting, visualization, new types of real-time analytics, and machine
learning.
A data warehouse is a database optimized to analyze relational data coming
from transactional systems and business applications. The data structure and schema
are defined in advance to optimize for fast queries. The data from a warehouse is
typically used for reporting and analysis. Data is cleaned, enriched, and transformed,
so it can act as the “single source of truth” that users can trust [23]. Many
organizations support both a data warehouse and a data lake, as they serve different
needs and use cases. A data lake enables storage of non-relational data from mobile
apps, IoT devices, and social media and does not require the schema to be defined
when data is captured (referred to as “schema on read”). Massive amounts of data
can be stored without careful schema design, thereby enabling flexibility in terms
of the kinds of questions or data analytics that might need to be performed in the
future. Data lakes enable different types of analytics including big data analytics,
text mining, real-time stream data processing, and machine learning.
One great, early example of a successful data lake is the Big Data implementation
at Mercy Hospital. The hospital leveraged technology to improve medical outcomes
for patients by utilizing one of the first comprehensive, integrated electronic health
record (EHR) systems to provide real-time, paperless access to patient information.
Utilizing the EHR from Epic Systems, every patient’s activity, including clinical
and financial interactions, was captured. The hospital needed to address several
typical challenges associated with Big Data implementations including scalability,
data schema requirements, and response to large data queries. Mercy, in partnership
with Hortonworks (one of the early Big Data providers at the time), has created
the Mercy Data Library, a Hadoop-based data lake. This data lake enabled the
integration, ingestion, and processing of large amounts of batch data extracts from
relational systems, real-time data directly from Epic HER, and information from
social media and even weather sources [24]. The combination of all of these data
sets in a common platform enables the hospital to ask and answer questions at that
were previously impossible due to scale, cost, or both.
One of the projects implemented on the Mercy system utilized thee advanced
analytics techniques on a large amount of intensive care unit (ICU) patients’ vitals
