A data science team’s tips to the ‘path of less wrongness’
The first principle is that you must not fool yourself, and you are the easiest person to fool. – Richard Feynman
Data scientists are having their day in the sun. The role is a strange stew of computer science, algorithm skills, hacking, statistical inferencing, probabilistic modeling, and scientific inquiry, and started off as a rebranding of statistical work. It became a bonafide job title only in 2008, yet it is now the “sexiest job of the 21st century.” However, amidst all the hoopla, there is a lot of confusion on what is required of a data scientist.
Although it is tempting to try and define a high-priesthood of the role, the fuzziness surrounding it is because different businesses require and therefore rightly organize themselves around different flavors of the role. However we choose to define it, becoming a data scientist requires a fundamental interdisciplinary mindset of problem-solving. The role of a data scientist at Mad Street Den (MSD) is centered around building data products using machine learning, statistical inferencing, and modeling of domain knowledge.
Here are some musings on our data science culture.
The soapbox comes before the speech
A reliable data platform is the cornerstone of an effective data science team. Our data platform instruments every product with usage data and pipelines and consolidates them strategically at different fidelities for use in data science workflows. Another key player in our data architecture is the careful deliberation of scale. Data comes in different shapes and sizes and the platform must accommodate all of them.
How I learned to stop worrying and love the error

Building a data-driven culture early on and embedding this into the DNA of the company is important for effective machine learning and data science implementation. Data science projects are characterized by weak contracts and experimental iterations defined by measures of accuracy, variability, and various tradeoffs. A process of scientific inquiry and decision-making must be followed to allow the algorithms to evolve.
All models are wrong, but some are useful
Now it would be very remarkable if any system existing in the real world could be exactly represented by any simple model. However, cunningly chosen parsimonious models often do provide remarkably useful approximations. For example, the law PV = RT relating pressure (P), volume (V), and temperature (T) of an “ideal” gas via a constant (R) is not exactly true for any real gas, but it frequently provides a useful approximation and, furthermore, its structure is informative since it springs from a physical view of the behavior of gas molecules. For such a model, there is no need to ask the question “Is the model true?” If “truth” is to be the “whole truth,” the answer must be “No.” The only question of interest is: “Is the model illuminating and useful?” – Box, 1978
Decades later, this is still relevant to data science. We begin by building an empirical model which makes generalizations about the environment it operates in based on evidence and prior knowledge, then we revise these beliefs under the right conditions.
Model knowledge, not information

In order to extract true signals from a firehose of various modalities of data (including visual, textual, behavioral, and others) coupled with human expertise, we must understand the underlying semantics of entities and their interconnections. Building abstractions of knowledge and connecting the causal dots is key to mining the patterns that will inform our data products.
Learning on the job
Avoiding filter bubbles with ‘informed serendipity’
Lessons on not picking up nickels in front of a steamroller
Seeing the forest through the trees
Can I get the icon in cornflower blue?
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.






