Tired of ads? Enjoy an ad-free experience by signing up.
  • Insights
    This article was written by a TIA community member. Insights pieces undergo the same rigorous editorial process that newsroom-produced articles have.
Lance Gutteridge · · 3 min read

People think big data is awesome, but it could just be bad math

In 2015 The Economist wrote an article about big data . Now, if any magazine should be numerate, I would think it would be The Economist (a magazine I subscribe to and am normally a great admirer of). However, in this case, it seems they failed to do the math and were taken in by the big data hype machine.

Let me explain why.

Math check

The article says that Walmart has collected 2.5 petabytes of data on customer transactions and does 1 million transactions an hour. It goes on to say that all of this data is stored in a massive data storage that is around 2.5 petabytes in size.

A petabyte is 1,000 terabytes or 1 million gigabytes.

Now, a million transactions an hour sounds like a lot. The article is fawningly awestruck by the sheer size of the data that is produced and stored.

But what is that data? Does the number of transactions really add up to that much data?

There are 8,760 hours in a year. At 1 million transactions an hour, that is 8.76 billion transactions a year. So, let’s go with 10 billion (even though that means a million transactions per hour for 24 hours a day, 365 days of the year, and then a major round up).

Now, what kind of data is in a transaction?

In the following, I’m going to assign a number of bytes for certain kinds of information. For example, standard storage for a date/time takes 8 bytes. Numbers like prices can be stored in 4 bytes. A credit card number is usually 16 digits but some can be up to 19. So, we will go with 19 bytes for a credit card.

In a transaction, someone goes into a Walmart store and buys some stuff. Let’s assume that the average number of items is five, which is high in my opinion.

In a transaction, there is header information that applies to the whole transaction and then there are items (we are assuming five).

Sometimes, this is a cash sale and not a lot is known about the customer. However, let us assume we have a date/time (8 bytes), a credit card (19 bytes), an affinity card (19 bytes), and the store ID — say 8 bytes. Information on the particular till taking the transaction, say 8 bytes. That is 62 bytes of header. Each item has a product code (12 bytes), a price (4 bytes), and a quantity (4 bytes). Which give us 20 bytes per item, which is 100 bytes for five items. That means that the transaction is around 162 bytes.

We may have missed some things here: there could be more than one product code on an item (external and internal) or there could be some discount codes for special offers. So to be safe, let’s blow up the size of the transaction. Instead of 162 bytes, we will say 1,000 bytes, which is ridiculously large for a transaction and should more than cover any extra information that we haven’t counted.

At our greatly inflated 10 billion transactions a year and our gross overestimation of 1,000 bytes per transaction, we have 10,000 billion bytes per year, which is 10 terabytes.

Let’s assume they are keeping 10 years of data. That would be 100 terabytes of data.

What’s going on?

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.

Community Writer

Lance Gutteridge

Dr. Lance Gutteridge has a Ph.D. in computability theory. He is currently CTO at Formever Inc. where he architects ERP authoring software.