Disconnect: no, machine translation is not about to break the language barrier
Disconnect is a weekly column in which Tech in Asia’s Charlie Custer pokes at holes, plays devil’s advocate, or otherwise attempts to rain on the tech industry’s parade.

Is the language barrier about to become a thing of the past? That’s the premise of this recent Wall Street Journal article, which argues that ten years from now we’ll have near-perfect speech-simultaneous machine translation that’ll even be able to recreate the speaker’s voice in your native language. This, author Alec Ross argues, will even bring the world closer together.
Mr. Ross is former senior adviser for innovation to the U.S. Secretary of State, so I certainly can’t compete with him on credentials. But I’m just going to say it anyway: Alec Ross is wrong.
Machine translation right now is pretty much garbage. Everyone knows this. For simple sentences in grammatically similar languages, it works OK. But give it a complex sentence to move between Chinese and English (for example), and the results are typically nonsense.
Ross argues that over the next ten years, the increasing mountain of language data machines have to process, and user-assisted translation input, will help improve the quality of machine translation significantly, and he’s not wrong about that. Machine translation is already pretty useful for simple stuff like “Where’s the bathroom?” and that will only get more true. But the kind of Star-Trek-style real-time perfect translation that Ross is describing is a lot more than ten years down the road. Here’s why:

It’s not as simple as just pressing the translate button.
Real-time translation is impossible. Well, unless your translation device can read minds, anyway. To understand why, you need to understand languages are structured differently. Take German, for example, in which the verb typically falls at the end of the sentence. To translate German into proper English, where verbs often come early in a sentence, a machine would need to hear the entire completed sentence in German before it could process the translation in English. That means that any in-your-ear translation device like Ross describes would have to be operating on at least a sentence-long delay.
Granted, accurate speech-to-speech with a one-sentence delay would still be pretty awesome. But sadly, we won’t even be getting that in ten years because…
Accents and dialects. These can cause big problems even for skilled human interpreters (and skilled human interpreters are much, much more accurate than machines, at least for now). If you’ve got a particularly thick accent and you’ve ever tried using Siri, you’ve probably experienced this firsthand. Then of course there’s also the matter of local dialects. And while machines could be programmed to learn dialects, getting them the raw data to work with will be more difficult. Take, for example, Dongbeihua a Mandarin dialect spoken in China’s northeast. It’s rarely if ever used in written communication, so all of the internet writing that machines could data mine is going to be pretty useless. Perhaps it would be possible if advanced machine learning programs had access to tons of recorded voice communications (for example, the WeChat logs of everyone in the Northeast), but even then the translation would be far from perfect because…
Cultural references. Language and culture are pretty inseparable. We don’t realize it, but our everyday speech is infused with references to aspects of our culture that can sometimes make it incomprehensible to outsiders, even if they understand the literal meaning of the words we’re saying. If I describe a task as Quixotic, for example, whether or not you know what that means depends very much on whether you grew up in a culture that’s familiar with Don Quixote. That sort of thing would be easy to program into a machine, of course, but language can move quickly and translating newer cultural references would be significantly harder.

This guy changed the meaning of “I believe” in Chinese overnight by saying something very dumb in public back in 2011.
For example, following the 2011 high speed rail crash in China, a government spokesperson gave a catastrophic press conference in which he used the phrase: “Whether or not you believe it, anyway I believe it.” The phrase went viral in China, and overnight people were mockingly saying “I believe it” when their actual meaning was the exact opposite of that. That’s the sort of thing a skilled human interpreter can catch pretty easily. But a machine is going to have a difficult time if the meaning of a seemingly innocuous phrase changes overnight due to some cultural factor. Maybe there could be a way to keep machine translation software up-to-date on the latest linguistic trends, but even then translations could still be misleading because…
Great machine translation isn’t happening without great AI
A dark ending
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.






