Tired of ads? Enjoy an ad-free experience by signing up.
Pradeep Menon · · 7 min read

Data science simplified: A breakdown of its key principles and processes

In 2006, UK mathematician and Tesco Clubcard architect Clive Humbly coined the phrase “Data is the new oil.” He said the following:

“Data is the new oil. It’s valuable, but if unrefined, it cannot be used. It has to be changed into gas, plastic, chemicals, etc. to create a valuable entity that drives profitable activity. So, data must be broken down, analyzed for it to have value.”

The iPhone revolution, growth of the mobile economy, and advancements in big data technology have created a perfect storm. In 2012, Harvard Business Review published an article that put data scientists on the radar. The article Data Scientist: The Sexiest Job of the 21st Century labeled this new breed of people as a hybrid of data hacker, analyst, communicator, and trusted advisor.

Every organization is now making attempts to be more data-driven. Machine learning techniques have helped them in this endeavor. I realize that a lot of the material out there is too technical and difficult to understand. In this series of articles, my aim is to simplify data science. I will take my cue from the Stanford book An Introduction to Statistical Learning.

Data science is a multi-disciplinary field. It is the intersection between the following domains:

  • Business knowledge
  • Statistical learning (aka machine learning)
  • Computer programming

In this article, I will begin by covering principles, general processes, and types of problems in the field.

Key principles

1. Data is a strategic asset

This concept is an organizational mindset. The questions to ask are “Are we using all the data assets that we are collecting and storing?” and “Are we able to extract meaningful insights from them?” I’m sure that the answers to these question are no. Cloud-born companies are intrinsically data-driven, as it is in their psyche to treat data as a strategic asset. This mindset is not valid for most of the organization.

2. A systematic process for knowledge extraction

A methodical process needs to be in place for extracting insights from data. This process should have clear and distinct stages with clear deliverables. The Cross Industry Standard Process for Data Mining (CRISP-DM) is one such process.

3. Sleeping with the data

Process

Machine learning problem types

Machine learning tasks to models to algorithms

Conclusion

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Pradeep Menon

Pradeep is an experienced Big Data and Data Science professional with 15+ years of experience. Pradeep works as a Cloud Solution Architect (CSA)- Advanced Analytics and AI with Microsoft.