- Premium Content It takes our newsroom weeks - if not months - to investigate and produce stories for our premium content. You can’t find them anywhere else.
Why Singapore’s LLM isn’t sweating GPT-4
What does the late Indonesian President Suharto, who ruled the country for 31 years, have to do with generative AI and the need for locally developed large language models (LLMs)?
Citing the long-deceased leader as an example, AI Singapore (AISG), a nonprofit organization connecting the country’s AI research institutions, startups, and companies, argues that Western-developed LLMs may harbor certain biases stemming from their training data and cultural context.
For instance, a version of Llama 2, the open-source LLM developed by Meta, portrayed Suharto largely negatively, emphasizing his role in political suppression and human rights abuses.
At a recent press event, AISG demonstrated how its model Sea-Lion, which stands for Southeast Asian Languages In One Network, also highlighted the achievements of Suharto, such as creating an environment of political stability and economic prosperity.

Photo credit: Shutterstock
This suggests Sea-Lion’s potential for handling nuanced perspectives on sensitive topics, particularly those requiring local context.
US-China AI arms race
One key argument driving organizations like AISG to develop their own LLMs is the need for models that are not solely aligned with Western culture.
Last month, Singapore announced its revised national AI strategy and committed S$70 million (US$52 million) to further develop Sea-Lion. However, some AI practitioners question the practicality and timing of the LLM project in light of OpenAI and other tech giants’ rapid advancements.
See also: Is Singapore’s LLM project timely or too late?
Still, amid the escalating US-China AI arms race, AISG believes there’s a need for diverse LLM alternatives.
Leslie Teo, a senior director at AISG, pointed out that perceived biases in AI models don’t stem from malicious intent or personal values of their creators. Instead, the origin and linguistic scope of training data play a crucial role in shaping model bias.
According to AISG, which collated data provided by AI tool developer Hugging Face, about 73% of existing LLMs originate from the US and China, and 95% of all models are primarily trained on English-language data or on a mix of English and one of Chinese, Arabic, or Japanese.
This means Southeast Asian languages like Bahasa Indonesian, Thai, and Tamil are severely underrepresented. As an example, AISG says that less than 0.5% of Llama 2’s training data comes from the region’s languages.
Sea-Lion aims to bridge this gap, claiming to be the first open-sourced LLM specifically focused on Southeast Asian languages and contexts. Around 64% of the data in the model is in English, while Southeast Asian languages comprise 13%, with the remainder taken up by Chinese and code.
Battle of the models
“Not trying to beat GPT-4”
Stay ahead in Asia’s tech landscape
This is premium content. Subscribe to read the full story.
AI Singapore believes that Sea-Lion’s ability to process Southeast Asian languages will complement existing models, including GPT-4.
We know this is not ideal. ⌛ Sign up in 20 seconds. Cancel anytime.
Our subscriber community includes professionals from these companies:





Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.

