Tired of ads? Enjoy an ad-free experience by signing up.
Janani Sriram · · 5 min read

A data science team’s tips to the ‘path of less wrongness’

The first principle is that you must not fool yourself, and you are the easiest person to fool. – Richard Feynman

Data scientists are having their day in the sun. The role is a strange stew of computer science, algorithm skills, hacking, statistical inferencing, probabilistic modeling, and scientific inquiry, and started off as a rebranding of statistical work. It became a bonafide job title only in 2008, yet it is now the “sexiest job of the 21st century.” However, amidst all the hoopla, there is a lot of confusion on what is required of a data scientist.

Although it is tempting to try and define a high-priesthood of the role, the fuzziness surrounding it is because different businesses require and therefore rightly organize themselves around different flavors of the role. However we choose to define it, becoming a data scientist requires a fundamental interdisciplinary mindset of problem-solving. The role of a data scientist at Mad Street Den (MSD) is centered around building data products using machine learning, statistical inferencing, and modeling of domain knowledge.

Here are some musings on our data science culture.

The soapbox comes before the speech

A reliable data platform is the cornerstone of an effective data science team. Our data platform instruments every product with usage data and pipelines and consolidates them strategically at different fidelities for use in data science workflows. Another key player in our data architecture is the careful deliberation of scale. Data comes in different shapes and sizes and the platform must accommodate all of them.

How I learned to stop worrying and love the error

Building a data-driven culture early on and embedding this into the DNA of the company is important for effective machine learning and data science implementation. Data science projects are characterized by weak contracts and experimental iterations defined by measures of accuracy, variability, and various tradeoffs. A process of scientific inquiry and decision-making must be followed to allow the algorithms to evolve.

All models are wrong, but some are useful

Now it would be very remarkable if any system existing in the real world could be exactly represented by any simple model. However, cunningly chosen parsimonious models often do provide remarkably useful approximations. For example, the law PV = RT relating pressure (P), volume (V), and temperature (T) of an “ideal” gas via a constant (R) is not exactly true for any real gas, but it frequently provides a useful approximation and, furthermore, its structure is informative since it springs from a physical view of the behavior of gas molecules. For such a model, there is no need to ask the question “Is the model true?” If “truth” is to be the “whole truth,” the answer must be “No.” The only question of interest is: “Is the model illuminating and useful?” – Box, 1978

Decades later, this is still relevant to data science. We begin by building an empirical model which makes generalizations about the environment it operates in based on evidence and prior knowledge, then we revise these beliefs under the right conditions.

Model knowledge, not information

In order to extract true signals from a firehose of various modalities of data (including visual, textual, behavioral, and others) coupled with human expertise, we must understand the underlying semantics of entities and their interconnections. Building abstractions of knowledge and connecting the causal dots is key to mining the patterns that will inform our data products.

Learning on the job

Avoiding filter bubbles with ‘informed serendipity’

Lessons on not picking up nickels in front of a steamroller

Seeing the forest through the trees

Can I get the icon in cornflower blue?


Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Janani Sriram

Janani is the head of the data science team at Mad Street Den, an Artificial Intelligence & Computer Vision startup.