214
F. Firouzi and B. Farahani
Table 4.6 The main differences between data lake and data warehouse
Data lake
Enterprise data warehouse (EDW)/RDBMS
Structured, semi-structured, unstructured data Structured, processed data
Physical collection of un-curated raw data
Data of common meaning
System of insight: unknown data to make
experimentation/data discovery
System of record: well-understood data to do
operational reporting
Any type of data
A limited set of data types (i.e., relational)
All workloads – batch, interactive, streaming,
machine learning
Optimized for interactive querying
Not suitable for transactions
Suitable for ACID transactions
Higher latency
Low latency (transactions)
Need skills to gain insights
Easily create reports (good support from BI
tools)
Expensive for large data
Low-cost data storage
Schema-on-read (ELT)
Schema-on-write (ETL)
new applications could be created to serve business needs. Currently, data lakes
support the following abilities:
• Cost-effective highly scalable storage of raw data
• Storage of diverse data types (structured, semi-structured, and unstructured data)
in the same place
• Query and transform data when a business question arises
• Defines data structure at the time of use (i.e., schema on reading)
• Incorporates new methods of data processing
As previously discussed, corporations have already started integrating data
lakes to address the requirements of their IoT projects. However, note that data
warehouses and data lakes were created to serve different user groups and different
purposes (see Table 4.6). In other words, data lake and data warehouse are complementary systems. Complementing a data warehouse with the addition of a data lake
is an agile step forward for the most companies. These combined solutions create
greater flexibility and improve speed when it comes to data processing, capturing
streaming, semi-structured, or unstructured data. It also provides increased data
warehouse bandwidth needed for business intelligence. Data lakes can be a powerful
tool useful for individuals, data scientists, and businesses desiring to prepare or
blend data or provide on-demand data profiling or to generate new insights from big
data. On the other hand, data warehouses are a better option for those who require
regularly published data that is already aggregated and processed. Finally note that
data warehouses solely work with structural data based on ETL approach, whereas
data lakes are low-cost storages to hold all types of data (structured and unstructured
data) working based on ELT approach.
F. Firouzi and B. Farahani
Table 4.6 The main differences between data lake and data warehouse
Data lake
Enterprise data warehouse (EDW)/RDBMS
Structured, semi-structured, unstructured data Structured, processed data
Physical collection of un-curated raw data
Data of common meaning
System of insight: unknown data to make
experimentation/data discovery
System of record: well-understood data to do
operational reporting
Any type of data
A limited set of data types (i.e., relational)
All workloads – batch, interactive, streaming,
machine learning
Optimized for interactive querying
Not suitable for transactions
Suitable for ACID transactions
Higher latency
Low latency (transactions)
Need skills to gain insights
Easily create reports (good support from BI
tools)
Expensive for large data
Low-cost data storage
Schema-on-read (ELT)
Schema-on-write (ETL)
new applications could be created to serve business needs. Currently, data lakes
support the following abilities:
• Cost-effective highly scalable storage of raw data
• Storage of diverse data types (structured, semi-structured, and unstructured data)
in the same place
• Query and transform data when a business question arises
• Defines data structure at the time of use (i.e., schema on reading)
• Incorporates new methods of data processing
As previously discussed, corporations have already started integrating data
lakes to address the requirements of their IoT projects. However, note that data
warehouses and data lakes were created to serve different user groups and different
purposes (see Table 4.6). In other words, data lake and data warehouse are complementary systems. Complementing a data warehouse with the addition of a data lake
is an agile step forward for the most companies. These combined solutions create
greater flexibility and improve speed when it comes to data processing, capturing
streaming, semi-structured, or unstructured data. It also provides increased data
warehouse bandwidth needed for business intelligence. Data lakes can be a powerful
tool useful for individuals, data scientists, and businesses desiring to prepare or
blend data or provide on-demand data profiling or to generate new insights from big
data. On the other hand, data warehouses are a better option for those who require
regularly published data that is already aggregated and processed. Finally note that
data warehouses solely work with structural data based on ETL approach, whereas
data lakes are low-cost storages to hold all types of data (structured and unstructured
data) working based on ELT approach.
