Data science simplified: A breakdown of its key principles and processes
In 2006, UK mathematician and Tesco Clubcard architect Clive Humbly coined the phrase “Data is the new oil.” He said the following:
“Data is the new oil. It’s valuable, but if unrefined, it cannot be used. It has to be changed into gas, plastic, chemicals, etc. to create a valuable entity that drives profitable activity. So, data must be broken down, analyzed for it to have value.”
The iPhone revolution, growth of the mobile economy, and advancements in big data technology have created a perfect storm. In 2012, Harvard Business Review published an article that put data scientists on the radar. The article Data Scientist: The Sexiest Job of the 21st Century labeled this new breed of people as a hybrid of data hacker, analyst, communicator, and trusted advisor.
Every organization is now making attempts to be more data-driven. Machine learning techniques have helped them in this endeavor. I realize that a lot of the material out there is too technical and difficult to understand. In this series of articles, my aim is to simplify data science. I will take my cue from the Stanford book An Introduction to Statistical Learning.
Data science is a multi-disciplinary field. It is the intersection between the following domains:
- Business knowledge
- Statistical learning (aka machine learning)
- Computer programming
In this article, I will begin by covering principles, general processes, and types of problems in the field.
Key principles

1. Data is a strategic asset
This concept is an organizational mindset. The questions to ask are “Are we using all the data assets that we are collecting and storing?” and “Are we able to extract meaningful insights from them?” I’m sure that the answers to these question are no. Cloud-born companies are intrinsically data-driven, as it is in their psyche to treat data as a strategic asset. This mindset is not valid for most of the organization.
2. A systematic process for knowledge extraction
A methodical process needs to be in place for extracting insights from data. This process should have clear and distinct stages with clear deliverables. The Cross Industry Standard Process for Data Mining (CRISP-DM) is one such process.
3. Sleeping with the data
Process
Machine learning problem types
Machine learning tasks to models to algorithms
Conclusion
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.






