WangchanX
Thailand's flagship open-source Thai LLM family — fine-tuning pipeline, instruction models, and benchmarks from VISTEC.
Overview
WangchanX is the umbrella Thai NLP project from the VISTEC-depa AI Research Institute of Thailand, spanning pre-trained encoders (WangchanBERTa), instruction-tuned generative models (WangchanLion, WangchanGLM), a model-agnostic fine-tuning pipeline (SFT, DPO, ORPO), human-annotated datasets (WangchanThaiInstruct — 75K examples across medical, legal, finance, and retail), and MRC evaluation tooling. All code and weights are released under Apache 2.0 on GitHub and Hugging Face, making it the most comprehensive open Thai language AI stack available.
Initiated by VISTEC's School of Information Science and Technology and co-funded by PTT, SCBX, and SCB, WangchanX has produced models that outperform OpenThaiGPT and SeaLLM on Thai machine reading comprehension F1 scores. The project collaborates with AI Singapore on the WangchanLION line (SEA-LION base, Thai instruction-tuned), positioning it as both Thailand's national AI language effort and a key node in the broader ASEAN open-LLM ecosystem. Deployment is self-hosted via TGI, LocalAI, or Ollama; no proprietary API or SaaS endpoint exists.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- Model weights on Hugging Face
- SFT / DPO / ORPO fine-tuning scripts on GitHub
- Thai instruct datasets (120k) included
- Commercial use permitted
- You pay only your own infrastructure costs
- No hosted API from AIResearch.in.th
Use cases
What you can produce with WangchanX
- A custom Thai-language instruction model fine-tuned on your own domain data using WangchanX SFT scripts
- A self-hosted Thai chatbot or Q&A system deployed via TGI or Ollama with no data leaving your premises
- A Thai legal document RAG pipeline answering questions about Thai Corporate and Commercial Law
- Benchmark scores for any Thai LLM on MRC tasks (XQuAD, iapp_wiki_qa_squad) using WangchanX MRC Eval
- A Thai sentiment classifier or NER tagger built on WangchanBERTa fine-tuned for your target domain
- A synthetic Thai instruction dataset generated using the seed-free pipeline (wangchanx-seed-free-synthetic-instruct-thai-120k methodology)
- A reproducible evaluation report comparing Thai LLMs using the WangchanThaiInstruct test split across medical, legal, finance, and retail domains
ASEAN Perspective
WangchanX in Southeast Asia
WangchanX is Thailand's most prominent nationally-funded open LLM initiative, developed inside the VISTEC-depa AI Research Institute — a government-backed (depa) research body. Its collaboration with AI Singapore on the WangchanLION / SEA-LION line makes it a concrete bridge between Thai and wider ASEAN NLP ecosystems, supporting the region's push for sovereign, locally-hosted AI. Because all weights are self-hostable under Apache 2.0, Thai organisations in regulated sectors (government, healthcare, finance) can run models entirely within Thai data-centre boundaries, satisfying PDPA data-residency expectations without reliance on US or Chinese cloud APIs. The WangchanX-Legal-ThaiCCL-RAG dataset also demonstrates domain relevance to ASEAN legal systems, which share some structural similarities with Thai corporate and commercial law.
WangchanX delivers exceptional value for Thai-language AI work: it is the only comprehensive, fully open-source Thai LLM stack with a coherent pipeline from pre-training data through instruction tuning, alignment (DPO/ORPO), and MRC evaluation. WangchanBERTa (100K+ monthly downloads) remains the dominant Thai encoder, and WangchanLion outperforms OpenThaiGPT and SeaLLM v2 on Thai F1 benchmarks. For Thai government, academic, and enterprise teams needing data-residency guarantees, nothing else matches its breadth.
The caveats are significant: there is no hosted API, no consumer chat interface, and setup requires ML engineering expertise (QLoRA, DeepSpeed, Flash Attention 2). GitHub activity is modest (46 stars, no formal releases), the largest generative models top out at 13B parameters, and the ecosystem lags commercial competitors like SCB 10X's Typhoon 2 on instruction-following quality and multimodal capability. Best suited for Thai-language researchers and enterprises willing to invest in self-hosted infrastructure.
What people say
Forty-seven GitHub stars is not a lot for a foundation-model project, and that's the honest starting point for evaluating WangchanX: this is academic infrastructure, not a product with a fanbase, and it doesn't try to be one. VISTEC's AI Research Institute has kept it going since WangchanBERTa's 2021 release, and the string of downstream work — WangchanGLM, WangchanLion, and the WangchanX fine-tuning pipeline with SFT/DPO/ORPO scripts — makes it the only actively maintained encoder-to-instruction-tuned stack built specifically for Thai.
Researchers treat WangchanBERTa (100K+ monthly HuggingFace downloads) as the default Thai encoder baseline, and the WangchanThaiInstruct dataset — 75,000 human-annotated examples across medical, legal, finance and retail domains — got an EMNLP 2025 paper out of it. WangchanLion, built jointly with AI Singapore off the SEA-LION base, beats OpenThaiGPT and SeaLLM v2 on Thai reading-comprehension F1 in VISTEC's own benchmarks — worth treating as a lower bound on the actual gap, given the source.
None of this ships as a product. There's no hosted API, no chat interface, and running any of it requires comfort with QLoRA, DeepSpeed or Flash Attention 2. Generative models cap out at 13B parameters, well short of what commercial Thai competitors like SCB 10X's Typhoon 2 now offer. For self-hosted, data-sovereign Thai NLP work, though, nothing else covers as much of the pipeline.
Summary of public user & expert reviews, compiled by RECATOOLS.
Notable facts
- The name 'Wangchan' comes from Wangchan Valley in Rayong — the science and technology zone where VISTEC's campus is physically located.
- WangchanBERTa was trained on 381 million unique Thai sentences (78.5 GB) using 8 NVIDIA V100 GPUs and has since spawned 58 community fine-tuned variants.
- WangchanLion is a joint Thai–Singaporean effort, built on AI Singapore's SEA-LION base model and then instruction-tuned specifically for Thai — demonstrating cross-border ASEAN AI collaboration.
- The WangchanThaiInstruct dataset (75K human-annotated examples, EMNLP 2025) is among the largest human-authored Thai instruction datasets ever released, covering medical, legal, finance, and retail domains.
Frequently asked questions
About this listing
This entry was compiled from publicly available data including WangchanX's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with WangchanX unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to WangchanX directly →
Spotted something out of date? Suggest an update →
Alternatives to WangchanX
More in LLMs & Chat