Tired of ads? Enjoy an ad-free experience by signing up.
  • Insights
    This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
suyash katyayani · · 5 min read

Startups need to do data now, so we made a serverless big data pipeline

plan

Photo credit: Pexels.

The following is based on my talk on Purplle’s serverless data pipeline at AWS Summit 2016 in India.

Everyone is talking about big data. There is hardly any conversation in the tech world that doesn’t include big data. It gives companies a competitive advantage over others, aids decision making, and adds amazing analytical prowess in organizations of all shapes and sizes. Therefore, it’s imperative that startups take control of their data and start treating it like a first-class citizen.

But it’s easier said than done! Whenever I meet young startup CTOs, I find that they are often bogged down by the cost, volumes, and processes that comes along with big data.

Common notions are:

  • “Its too early for us.”
  • “We don’t have the varied skill sets in the team, which are required to manage this.”
  • “We are not ready for a huge spike in operating expenditure without an immediate ROI.”
  • “We want to finalize our big data strategy before going ahead with implementation.”

My suggestion: Start doing data now.

Be it events, logs, or click streams, the strategy of creating value from it can keep on evolving. We are at a point in time where technology has progressed that it no longer takes a rocket scientist to extract information and derive actionable insights. There are amazing visualization tools and database warehouses like Redshift and Big Query that enable people from various backgrounds to become data masters. Having said that, you need to ensure that there are minimal overheads in terms of people, operating costs, and infrastructure management.

We at Purplle.com have created a scalable, serverless data pipeline  at low operating costs and engineered by only a single developer. It helps us collect millions of data points every day.

The thought process for our data pipeline

Let’s start from the beginning . We sat down and sketched out a basic architecture of the data flow from event producers to data lake.

architecture

Big data pipeline architecture.

Challenges

  • Variety: Diverse data sources  (apps, web, CRM) and formats  (events, chats, structured data, unstructured text)
  • Velocity : Uneven capacity needs that are typically millions a day and split into crests, troughs, and spikes
  • Veracity : Biases, noise, and abnormalities in data

We tried to define the ideal infrastructure needed for each leg of the data pipeline.

Solution : Think serverless


Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

suyash katyayani

Quintessential startup techie. The CTO Of Purplle.com. Has great passion for engineering products and experimenting with new tech. An IITKgp graduate, Foodie, Coffee Lover, Wanna-be Singer