Is China’s data an unbeatable AI advantage? Not really, says report
In this dawning age of AI, the conventional wisdom is that data is the new oil and China is the new Organization of the Petroleum Exporting Countries.
But a new report released on Tuesday suggests that the staggering amount of data generated by China’s 1.4 billion population may not be as big an advantage in the global AI competition as it was thought to be.

Photo credit: Gavin Tracy
The report by MacroPolo, the in-house think tank of the Paulson Institute in Chicago, argues that data is not a single-dimensional resource for AI. And despite China’s formidable data reserves, the US still holds advantages in data quality and diversity.
“Many assume the size of China’s population gives it an advantage in the volume of data, but this is actually misleading,” said Sheehan, a San Francisco-based fellow at MacroPolo who wrote the paper.
“The relationship between data and AI prowess is analogous to the relationship between labor and the economy. China may have an abundance of workers, but the quality, structure, and mobility of that labor force is just as important to economic development,” Sheehan added.
The research comes as the US and China, competing on so many economic and cultural fronts, are locked in a rivalry over AI technologies.
In 2017, China’s State Council issued a three-step plan to make the country a global leader in AI by 2030. Just this February, US President Donald Trump issued an executive order intended to maintain America’s global AI leadership, directing government agencies to prioritize AI in their research and development spending.
To a significant degree, the competition in AI is a competition for data. From facial recognition to autonomous cars and machine translation, most AI applications are possible only after machines study a huge amount of data and find hidden patterns between inputs and outcomes. Only then could a machine start to learn how to master human skills.
Thus, data today is regarded by many technologists as an important – if not vital – strategic resource for the AI economy.
But in his MacroPolo paper, Sheehan broke down data into five different dimensions: quantity, depth, quality, diversity, and access. The paper, more of a framework than a quantitative research, found that China and the US are tied in the quantity of their data. China holds advantages in terms of data depth and access, while the US has superior data quality and diversity.
More than 800 million Chinese have connected to the internet, generating abundant data concerning a wide array of online activities – from grocery shopping to buying wealth-management products and booking a table at a restaurant.
But most internet service providers in China still largely focus on their domestic market, while Silicon Valley companies have a more global reach. Users of Google and Facebook represent a far greater range of languages, ethnicities, cultures, and nationalities than those of WeChat, China’s No. 1 social messaging tool, whose 1 billion users are almost all Chinese.
As a result, an AI-operated facial recognition program, for example, may have difficulty identifying people other than Chinese if all the data it has studied is exclusively of Chinese faces.
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.







