Taiwan’s sovereign AI paradox: What exactly do we mean by sovereign?
As many countries reconsider dependence on US and Chinese AI companies, Taiwan is prioritizing local control over narratives about the country such as democracy and the rule of law
By Nigel P. Daly / Contributing reporter
Taiwan wants AI to know more about Taiwan. On Sept. 15, the Ministry of Digital Affairs (MODA) invited publishers, writers and electronic-book platforms to contribute to Taiwan’s Sovereign AI Training Corpus. The database now contains about 5,000 datasets and well over a 2.2 billion tokens of text.
Digital Minister Lin Yi-jing (林宜敬) explained that an AI model’s understanding of ideas such as democracy and checks and balances is shaped by its training data.
To let AI better understand Taiwan, he said, Taiwan’s language, culture and values need to become part of what AI learns.
A message reading “AI artificial intelligence,” a keyboard and robot hands are seen in this illustration taken in January last year.
Photo: Reuters
Taiwan is not alone in its move toward “sovereign AI.” Many countries are reconsidering dependence on US and Chinese technology companies. Chinese models can carry political assumptions shaped by the Chinese Communist Party, while American companies control access, prices and development of the most capable frontier systems.
Here, sovereign AI has three dimensions: Knowledge sovereignty means AI understands Taiwan accurately; Data sovereignty means sensitive information remains under local control; and Technological sovereignty means Taiwan is not overly dependent on a foreign model provider. Switzerland’s open Apertus model similarly prioritizes transparency and local control rather than competing with frontier models.
WHEN CHINESE DID NOT MEAN TAIWANESE
Taiwan launched the government-backed Trustworthy AI Dialogue Engine, or TAIDE, in 2023 to better handle Traditional Chinese and Taiwanese knowledge.
The need was clear. At the time, TAIDE project leader Lee Yu-chieh (李育杰) noted that BLOOM, a 2022 multilingual model from the international BigScience project, had trained on 16 percent Simplified Chinese data but only 0.05 percent Traditional Chinese.
Shortly before National Day in 2023, an Academia Sinica Traditional Chinese model based partly on Simplified Chinese datasets also gave incorrect answers about Taiwan’s National Day and president.
With Oct. 10 approaching, the lesson remains relevant.
TAIDE addressed this by adapting existing open-weight models, whose trained parameters can be downloaded and modified. Early versions used Meta’s Llama, while the current Gemma-3-TAIDE-12B is based on Google’s Gemma 3 12B, released in March last year.
Even in 2023, iKala founder Sega Cheng (程世嘉) questioned the economics of building a national foundation model from scratch.
“Anyone with some sense of the costs would not rush headlong into training their own foundation model,” he said. TAIDE’s later development largely followed that logic: Sovereign AI 1.0 meant taking a capable foreign base model and teaching it more about Taiwan.
DOES LOCALIZATION HELP?
Quite a bit, but only up to a point. MODA’s sovereign corpus contains 2.2 billion tokens, the units of text processed by an AI. Google trained Gemma 3 12B on about 12 trillion tokens. The scales are different because the datasets serve different purposes: foundation models need vast, diverse data, while Taiwan’s corpus adds concentrated local knowledge through further training, fine-tuning or retrieval.
Tests show that this localization works. TMMLU+, a Taiwanese benchmark developed by researchers including Cheng, contains more than 20,000 multiple-choice questions across 66 subjects reflecting Taiwanese contexts. The February TAIDE model scored 58 percent overall, compared with 54 percent for the original Gemma model. On Taiwan’s geography, it scored 70 percent versus 61 percent.
That is a clear improvement, but 58 percent still means getting roughly four questions in 10 wrong. Meanwhile, on the current TMMLU+ v1.1 leaderboard, Claude Opus 5 scores 96 percent, Grok 4.3 scores 76 percent and Google’s open-weight Gemma 4 31B scores 70 percent. The benchmark has since been revised, so these scores are not directly comparable, but they still illustrate a wide capability gap.
Taiwan-focused training improves local knowledge but does not erase that gap. If Taiwan closes it by relying on proprietary foreign frontier models, it gains capability at the cost of technological sovereignty.
FROM SOVEREIGN MODEL TO SOVEREIGN SYSTEM
That compromise helps explain why Taiwan’s strategy has expanded beyond TAIDE. MODA’s corpus can also support retrieval-augmented generation (RAG). Instead of trying to put every fact inside a model during training, RAG searches a trusted database and passes relevant documents to the AI before it answers.
Still, while RAG can improve accuracy, it can’t guarantee it. An April benchmark study found that while standard RAG provided increasing gains on complex queries, combining retrieval with dedicated large reasoning models improved accuracy by an additional 10 percent in that benchmark.
“A stronger reasoning model grounded in the same trusted Taiwanese data,” says Yung-Chun (Peter) Chang (張詠淳), Deputy Director of Taipei Medical University’s Research Center for AI in Medicine, “could outperform a weaker locally adapted model,” especially for tasks requiring “complex interpretation or synthesis across documents.”
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.