Pleias

Small open LLMs trained only on permissively licensed data

LLMs & Chat Open Source Open Source
Researched · Published · Reviewed
RECATOOLS Score
6 / 10
Capability
5
Value for money
8
Ease of use
5
ASEAN readiness
6
API quality
4
Founded
HQ
Users
Launched
Developer

Overview

Pleias is a Paris lab building small (350M-3B parameter) open models trained only on Common Corpus, its ~2-trillion-token dataset of permissively licensed text. Its Pleias-RAG models target retrieval-heavy, regulated use cases where data provenance and EU AI Act compliance matter more than scale.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

What you can produce with Pleias

  • Common Corpus: ~2-trillion-token, fully open multilingual pretraining dataset
  • Pleias 1.0 model family (350M/1.2B/3B) across 8 EU languages
  • Pleias-RAG-350m and -1B, citation-grounded retrieval models
  • CommonLingua language-ID model covering 334 languages (with GSMA)
  • Nemotron-Personas-France and -Belgium synthetic datasets (with NVIDIA)
  • Open-source training/inference code on GitHub (Pleias-RAG-Library, nanotron-pleias)
Advertisement

ASEAN Perspective

Pleias in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Pleias takes the opposite bet from the frontier labs: instead of scraping everything and scaling up, it trains small (350M to 3B parameter) models exclusively on Common Corpus, its own roughly 2-trillion-token dataset of public-domain and permissively licensed text — accepted as an oral presentation at ICLR 2026. That data discipline is the actual product for regulated industries worried about copyright exposure under the EU AI Act, more than raw benchmark performance.

The RAG-specialized models (Pleias-RAG-350m, -1B) hold their own against larger 4B-8B rivals like Qwen2.5 and Llama-3.1 on retrieval benchmarks, and ship with built-in source citation. There's no hosted commercial product or API — everything ships as open weights on Hugging Face and GitHub, so this is a self-host, developer-only option. Recent moves (an NVIDIA persona-dataset collaboration, a 334-language identification model with GSMA) suggest a lab expanding its open-data footprint rather than chasing consumer traction.

Independent AI-assisted assessment by RECATOOLS.

What people say

There's no G2 page or app-store rating for Pleias — it's a research lab shipping open weights and datasets, not a SaaS product, so the usual review-site signal doesn't apply. What stands in for it is academic and developer reception, and that's been substantial for a company founded only in 2023.

Common Corpus, Pleias' roughly 2-trillion-token pretraining dataset built entirely from public-domain and permissively licensed sources — literature, government and legal documents, scientific papers, code — was accepted as an oral presentation at ICLR 2026, a real signal of technical credibility in a field where most "open" datasets still have murky provenance. Pleias 1.0, the model family trained on it, comes in 350M, 1.2B and 3B parameter sizes, instruction-tuned across eight European languages, and is described in its own paper as the first family trained exclusively on fully open data at that scale.

The more practical release is Pleias-RAG: 350M and 1B models mid-trained specifically for retrieval-augmented generation, with citation-grounded, "proto-agentic" behavior — they assess a query's complexity and decide how to respond based on whether the retrieved sources actually support an answer. On HotPotQA and 2WikiMultihopQA benchmarks, Pleias reports these small models are competitive with much larger ones, including Qwen-2.5-7B, Llama-3.1-8B and Gemma-3-4B, which is the headline claim worth stress-testing yourself given it comes from the lab's own paper.

2026 has brought two notable expansions: CommonLingua, a language-identification model covering 334 languages including 61 African languages, built with the GSMA; and Nemotron-Personas-France and -Belgium, synthetic population datasets built with NVIDIA to let regulated industries like banking and healthcare simulate realistic documents without using real personal data.

On GitHub, the Pleias org carries about 23 repositories, including Pleias-RAG-Library and nanotron-pleias (a minimalist 3D-parallel training framework), all actively maintained. There's no commercial pricing because there's no hosted product to buy — everything here is meant to be downloaded, fine-tuned or self-hosted, which narrows the audience to developers and researchers rather than end users.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Pleias's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Pleias unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Pleias directly →

Spotted something out of date? Suggest an update →

Advertisement