Tired of ads? Enjoy an ad-free experience by signing up.
  • Insights
    This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
Pradeep Menon · · 9 min read

Demystifying the Data Lake architecture

According to Gartner, 80 percent of successful chief data officers will have value creation or revenue generation as their first priority through 2021.

New architectural patterns need to be developed to harness the power of data. To fully capture the value of using big data, organizations need to have flexible data architectures and be able to extract maximum value from their data ecosystem.

The Data Lake concept has been around for some time now. However, I have seen organizations struggle to understand the concept as many of them are still boxed in the older paradigm of Enterprise Data Warehouses (EDW).

In this article, I will dive deep into the conceptual constructs of the Data Lake architecture pattern and lay out an architecture pattern.

Let us start with the known.

The traditional Enterprise Data Warehouse architecture

The traditional EDW architecture pattern has been used for many years. There are data sources which are extracted, transformed, and loaded (ETL). Then, along the way, we do some kind of structure creation, cleansing, etc. We pre-define the data model in EDW (dimensional model or 3NF model) and then create departmental data marts for reporting, OLAP (online analytical processing) cubes for slicing and dicing, and self-service BI.

This pattern is quite ubiquitous and has served us well for a long time now.

However, there are some inherent challenges in this pattern that can’t scale in the era of big data:

  • Our philosophy is that we need to understand the data first. What is the source system structure? What kind of data does it hold? What is the cardinality? How should we model it based on the business requirements? Are there any anomalies in data? And so forth. This is tedious and complex work. I used to spend at least two to three months in the requirement and data analysis phase. The EDW projects span for a few months to a few years. And this is all based on the assumption that the business knows the requirements.
  • We also have to make choices and compromises on which data to store and which data to discard. A lot of time is spent upfront on deciding what to bring in, how to bring them in, how to store and transform them, etc. Lesser time is spent on actually performing data discovery, uncovering patterns, or creating a new hypothesis for business value-add.

Definition of data in the world of big data

Let us now discuss briefly how the definition of data has changed. The four Vs of big data are now well-known—volume, velocity, variety, and veracity. Let me put some context to these things:

  • Data volumes have exploded since the iPhone revolution. There are 6 billion smartphones and nearly 1 PB of data is created every year.
  • Data is not just at rest. There are streaming data, IoT devices, and a plethora of data emanating from multiple fronts.
  • It is also about the variety of data. Video feeds and photographs are all data points now that demand to be analyzed and exploited.
  • With the explosion of data also comes the challenge of data quality. Which one should be trusted and which one should not be is a bigger challenge in the big data world.

Moore’s Law

The Data Lake analogy

Conceptual Data Lake architecture

Key differences

Data Lake on Azure

Key takeaways

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Pradeep Menon

Pradeep is an experienced Big Data and Data Science professional with 15+ years of experience. Pradeep works as a Cloud Solution Architect (CSA)- Advanced Analytics and AI with Microsoft.