4 Architecting IoT Cloud
213
4.6.3 Data Lake
The data lake has been created to address big data and the shortcomings of
traditional databases and data warehouses (see Fig. 4.28). James Dixon, the founder
and CTO of Pentaho, defined data lakes in the following way: “If you think of a
traditional relational database as a store of bottled water – cleansed and packaged
and structured for easy consumption – the data lake is a large body of water in a
more natural state. The contents of the data lake stream in from a source to fill the
lake, and various users of the lake can come to examine, dive in, or take samples.”
In contrast to DWH which only holds cleaned transformed structured data, a data
lake is a data-centered architecture that stores a vast amount of both structured and
unstructured data in its raw format until it is needed. Indeed, a key advantage of
data lakes is their capability to cost-effectively store data of unknown importance
or value that would generally be removed because of the cost required to store
the data securely. As a business’ analytic abilities grow, the prospective data use
cases are revealed, enabling historical data to be utilized in the training of machine
learning models or answer future questions. Note that data lakes store the raw data,
thereby, when they receive a business question, they need to query and transform
data to be able to address the question. The data lake can be built using many
technologies such as Amazon Simple Storage Service, Hadoop, NoSQL DBs, or
various combinations to support a variety of formats such as images, logs, Excel,
CSV, and sensor data. It has been discovered that as more data became available,
Fig. 4.28 The data lake has been designed to address big data
Précédent

- 220/647

Suivant