216
F. Firouzi and B. Farahani
4.6.3.2 Challenges of Data Lakes
While data lakes are a wonderful data management solution in our data-driven
environment, they can come with challenges that are worthy of attention. Gartner
has reminded, “the data lake will end up being a collection of disconnected data
pools or information silos all in one place. Without descriptive metadata and a
mechanism to maintain it, the data lake risks turning into a data swamp.”
To address such risks, a governing layer must be added to the architecture to
answer the questions below and ensure the four pillars of data governance are
included. The data governance is a union of tooling and processes that usually
increases the solution’s total cost of ownership (TCO), but on the other hand, it
increases the chance of return on investment (ROI) as well. The governance in data
lakes must be created with the following questions and elements in mind [27]:
• Data Catalog (What is the data and where is the data stored?) – Data catalogs
enable the user to compile and review metadata and index data so that they
become searchable. Metadata is useful for auditing or for actively driving data
transformation. Some data catalogs are also able to manually tag data to note that
it includes personally identifiable information (PII) or utilize a machine learning
algorithm to identify sensitive data.
• Data Quality (Is the data accurate and useful for a specific purpose?) –
Achieving data quality can be done using master data management (MDM),
a foundational process utilized to synchronize, categorize, centralize, organize,
enrich, or localize master data based on a business’ operational strategy. However, it is important to remember that MDM is originally established in relational
data warehouses and databases, so it is compatible with structured data only. Best
practices would indicate that MDM be applied selectively within data lakes.
• Data Lineage (Where did the data come from and how has it been transformed?) – Data lineage refers to the origin of data, what is done to it, and
where it goes over time. Possessing the complete audit trail of data, including
origin data, transformation information, and its analytic uses, is a necessity for
meeting the regulatory requirements most organizations currently face. Data
lineage information can also assist engineers in debugging or troubleshooting
issues that arise while handling workloads.
• Data Security (Is the data safe from unauthorized access?) – When creating a
data lake, it is important to address the following data security issues:
– Role-based access control (at suitable granularity level)
– Network isolation (e.g., security groups, firewall rules)
– End-to-end encryption (e.g., SSL certificates)
F. Firouzi and B. Farahani
4.6.3.2 Challenges of Data Lakes
While data lakes are a wonderful data management solution in our data-driven
environment, they can come with challenges that are worthy of attention. Gartner
has reminded, “the data lake will end up being a collection of disconnected data
pools or information silos all in one place. Without descriptive metadata and a
mechanism to maintain it, the data lake risks turning into a data swamp.”
To address such risks, a governing layer must be added to the architecture to
answer the questions below and ensure the four pillars of data governance are
included. The data governance is a union of tooling and processes that usually
increases the solution’s total cost of ownership (TCO), but on the other hand, it
increases the chance of return on investment (ROI) as well. The governance in data
lakes must be created with the following questions and elements in mind [27]:
• Data Catalog (What is the data and where is the data stored?) – Data catalogs
enable the user to compile and review metadata and index data so that they
become searchable. Metadata is useful for auditing or for actively driving data
transformation. Some data catalogs are also able to manually tag data to note that
it includes personally identifiable information (PII) or utilize a machine learning
algorithm to identify sensitive data.
• Data Quality (Is the data accurate and useful for a specific purpose?) –
Achieving data quality can be done using master data management (MDM),
a foundational process utilized to synchronize, categorize, centralize, organize,
enrich, or localize master data based on a business’ operational strategy. However, it is important to remember that MDM is originally established in relational
data warehouses and databases, so it is compatible with structured data only. Best
practices would indicate that MDM be applied selectively within data lakes.
• Data Lineage (Where did the data come from and how has it been transformed?) – Data lineage refers to the origin of data, what is done to it, and
where it goes over time. Possessing the complete audit trail of data, including
origin data, transformation information, and its analytic uses, is a necessity for
meeting the regulatory requirements most organizations currently face. Data
lineage information can also assist engineers in debugging or troubleshooting
issues that arise while handling workloads.
• Data Security (Is the data safe from unauthorized access?) – When creating a
data lake, it is important to address the following data security issues:
– Role-based access control (at suitable granularity level)
– Network isolation (e.g., security groups, firewall rules)
– End-to-end encryption (e.g., SSL certificates)
