- Insights This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
Startups need to do data now, so we made a serverless big data pipeline

Photo credit: Pexels.
The following is based on my talk on Purplle’s serverless data pipeline at AWS Summit 2016 in India.
Everyone is talking about big data. There is hardly any conversation in the tech world that doesn’t include big data. It gives companies a competitive advantage over others, aids decision making, and adds amazing analytical prowess in organizations of all shapes and sizes. Therefore, it’s imperative that startups take control of their data and start treating it like a first-class citizen.
But it’s easier said than done! Whenever I meet young startup CTOs, I find that they are often bogged down by the cost, volumes, and processes that comes along with big data.
Common notions are:
- “Its too early for us.”
- “We don’t have the varied skill sets in the team, which are required to manage this.”
- “We are not ready for a huge spike in operating expenditure without an immediate ROI.”
- “We want to finalize our big data strategy before going ahead with implementation.”
My suggestion: Start doing data now.
Be it events, logs, or click streams, the strategy of creating value from it can keep on evolving. We are at a point in time where technology has progressed that it no longer takes a rocket scientist to extract information and derive actionable insights. There are amazing visualization tools and database warehouses like Redshift and Big Query that enable people from various backgrounds to become data masters. Having said that, you need to ensure that there are minimal overheads in terms of people, operating costs, and infrastructure management.
We at Purplle.com have created a scalable, serverless data pipeline at low operating costs and engineered by only a single developer. It helps us collect millions of data points every day.
The thought process for our data pipeline
Let’s start from the beginning . We sat down and sketched out a basic architecture of the data flow from event producers to data lake.

Big data pipeline architecture.
Challenges
- Variety: Diverse data sources (apps, web, CRM) and formats (events, chats, structured data, unstructured text)
- Velocity : Uneven capacity needs that are typically millions a day and split into crests, troughs, and spikes
- Veracity : Biases, noise, and abnormalities in data
We tried to define the ideal infrastructure needed for each leg of the data pipeline.
Solution : Think serverless
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.







