AI Singapore, Google partner to improve large language models in SEA

Google Asia Pacific office opening ceremony in 2016 with the attendance of Prime Minister Lee Hsien Loong / Photo credit: Google
AI Singapore (AISG) and Google Research are collaborating to improve datasets for training and evaluating large language models (LLMs) in Southeast Asian languages.
Under Project SEALD, which stands for Southeast Asian Languages in One Network Data, the pair seeks to enhance the capabilities of such models, making them more useful across the region for the benefit of society.
Beginning with Indonesian, Thai, Tamil, Filipino, and Burmese, Project SEALD wants to create rich and diverse datasets for languages spoken in Southeast Asia.
One example involves the use of LLMs to bridge communication gaps with underrepresented migrant worker communities in Singapore, where individuals may be more fluent in regional languages than in English. Training LLMs with more culturally relevant data will help foster stronger engagement between the Singapore government, employers, and the migrant worker population.
“This [project] will open new opportunities and make AI more inclusive, accessible, and helpful for individuals and businesses throughout the region,” Yolyn Ang, vice president of Asia-Pacific business development at Google.
Additionally, SEALD will contribute to the development of AISG’s model Sea-Lion (Southeast Asian Languages in One Network). AISG and Google also plan to make the datasets and findings from Project SEALD available to the public through open-source channels.
Meanwhile, Google also has a similar partnership in India called Project Vaani, which focuses on the South Asian country’s diverse speech data across its 773 districts.
Editing by Miguel Cordon and Dhania Putri Sarahtika
(And yes, we’re serious about ethics and transparency. More information here.)
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.







