
KUALA LUMPUR (Sept 23): YTL AI Labs has released a dataset of 1.35 million synthetic Malaysian personas, developed in collaboration with Nvidia to help developers build and test artificial intelligence (AI) applications for local users.
Called Nemotron-Personas-Malaysia, the dataset was generated from 150,000 base records using probability distributions drawn from published Malaysian demographic statistics, YTL AI Labs said in a statement on Wednesday. It is available on Hugging Face under a CC BY 4.0 licence, which permits commercial use with attribution.
The company said the dataset contains no personal data and cannot be used to identify individuals.
Its 39 fields cover characteristics including age, gender, occupation, location and personality traits. According to YTL AI Labs, the personas are designed to reflect demographic patterns down to district level, including ethnic and regional differences across Peninsular Malaysia, Sabah and Sarawak.
YTL AI Labs CEO Foong Chee Mun said Malaysian languages, cultures and ways of working needed to be represented in the data used to develop AI applications.
The company said developers could use the personas to generate synthetic training data and test whether AI applications perform differently for various groups of users. It cited customer service, banking and government digital services as potential applications.
Nemotron-Personas-Malaysia is compatible with Nvidia NeMo libraries and is the first in Nvidia’s Nemotron-Personas collection to lead with Bahasa Melayu, according to YTL AI Labs.
..........
EdgeProp monthly brings you data, insights and solutions for an evolving market. Subscribe now for your free copy!
