YTL AI Labs, Nvidia launch 1.35mil persona Malaysian AI dataset

KUALA LUMPUR: YTL AI Labs has teamed up with Nvidia to launch an open dataset containing 1.35 million synthetic Malaysian personas, in a move aimed at giving developers and enterprises a locally grounded foundation for building and testing artificial intelligence (AI) systems.
Called Nemotron-Personas-Malaysia, the dataset was generated from 150,000 base records using demographic distributions derived from official Malaysian statistics.
In a statement, YTL AI Labs said the dataset is designed to capture Malaysia's demographic, linguistic, cultural and regional diversity, allowing AI developers to build systems that better reflect the population they are intended to serve.
The dataset is available free for commercial use under a Creative Commons Attribution 4.0 (CC-BY-4.0) licence and contains no personal data, with no individual identifiable from the synthetic personas.
The launch also strengthens YTL Group's push to build a broader sovereign AI ecosystem in Malaysia, spanning computing infrastructure, locally developed AI models and data designed around Malaysian users.
"The next generation of AI will not be defined only by who builds the largest model; it will be defined by how well intelligence understands the people it serves," YTL AI Labs chief executive officer Foong Chee Mun said.
"Malaysia has its own languages, cultures, institutions and ways of working. If AI is going to become part of our everyday lives, we need to make sure that Malaysian context is represented from the ground up."
The dataset is the latest addition to Nvidia's growing Nemotron-Personas collection, which includes datasets developed for markets including the United States, Japan, India, Singapore, Brazil, France, Korea, El Salvador, Vietnam and Belgium.
YTL AI Labs said the Malaysian dataset is the first in the collection to lead with Bahasa Melayu.
The collaboration combines Nvidia's synthetic data technology and methodology, including its NeMo Data Designer library and Nemotron-Personas framework, with YTL AI Labs' work to incorporate Malaysian demographic statistics and local linguistic, cultural and regional context.
The dataset can be used across the AI development cycle, including synthetic data generation, model fine-tuning, alignment, red-teaming and application testing.
For businesses, this could allow AI systems to be tested against different combinations of age, location, occupation and language before being deployed to customers.
A bank, for example, could use the synthetic personas to assess whether its AI system can handle different ways Malaysians communicate financial needs, while companies developing customer service applications could test performance across different demographic and regional groups.
Government agencies and researchers could similarly use the dataset to identify potential gaps in AI performance across different population groups.
The initiative forms part of YTL Group's broader sovereign AI strategy, which brings together computing infrastructure, AI models and locally relevant data.
YTL AI Cloud provides the computing infrastructure to train and run AI models in Malaysia, while YTL AI Labs develops locally focused technologies, including its ILMU family of models.
The company said the new dataset adds another layer to that ecosystem by providing developers with a structured representation of Malaysian users that can be used to build and evaluate AI applications.
The goal is to enable Malaysians to build intelligence for themselves, using world-class technology while retaining control over the context and priorities that make it uniquely Malaysian, it said.
Nemotron-Personas-Malaysia is available on Hugging Face under the CC-BY-4.0 licence and is compatible with Nvidia NeMo libraries.
It adds to YTL AI Labs' portfolio of Malaysian AI products, including the ILMU family of models, ILMU Claw for building autonomous AI agents through OpenClaw, and ILMUchat, its conversational AI assistant available through Apple's App Store and Google Play.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.