Tülu by AllenAI

Ai2's open playbook for turning base LLMs into instruct models

LLMs & Chat Open Source Open Source
Researched · Published · Reviewed
RECATOOLS Score
6.6 / 10
Capability
6
Value for money
8
Ease of use
5
ASEAN readiness
6
API quality
5
Founded
2023
HQ
Seattle, Washington, USA
Users
Launched
Developer

Overview

Tülu is Ai2's family of openly post-trained language models, built on Llama 3.1 and OLMo 2 bases and released with the full training data, recipes and eval framework used to make them — for researchers, not chatbot users.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

Use cases

Open RLHF research Self-hosted assistants Fine-tuning starting point

What you can produce with Tülu by AllenAI

  • Full SFT, DPO and RLVR training recipe published
  • 8B, 70B and 405B parameter checkpoints
  • Both Llama 3.1 and OLMo 2 base variants
  • Open-instruct training codebase and eval suite on GitHub
  • Training data and preference datasets released publicly
  • OLMo 2 variant available under Apache 2.0
Advertisement

ASEAN Perspective

Tülu by AllenAI in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Tülu is less a product than a paper trail: Ai2 publishes the SFT data, DPO preference pairs and the RLVR reinforcement method behind each release, so anyone can trace exactly how a base Llama 3.1 or OLMo 2 model became an instruction-follower. The 405B version, which followed the original Tülu 3 release, beat DeepSeek V3 on several benchmarks and traded blows with GPT-4o-mini — a real result for a fully open recipe. Since then Ai2 has folded the Tülu post-training methods into OLMo 3 (November 2025) rather than shipping a standalone Tülu 4, so this is best read as a milestone whose ideas live on elsewhere, not an actively iterating product line. There's no chat app or hosted API — you bring your own GPUs or a third-party host. Good for post-training researchers and ML engineers; not for anyone wanting a finished assistant.

Independent AI-assisted assessment by RECATOOLS.

What people say

Tülu 3's headline claim held up under scrutiny: the 70B checkpoint beat the instruct versions of Llama 3.1, Qwen 2.5 and Mistral at comparable sizes, and the 405B variant landed competitively against GPT-4o-mini and Claude 3.5 Haiku while beating DeepSeek V3 on a chunk of Ai2's eval suite. VentureBeat and other outlets covered it as proof that open post-training recipes could close the gap with closed labs' fine-tuning — not just match base-model benchmarks, which open models had already done, but match the harder-to-replicate alignment step.

The reaction on Hacker News was more measured than the press coverage. Commenters on the 405B thread noted that Ai2's contribution was the recipe, not the scale — nobody at Ai2 threw DeepSeek-level compute at it, and the interesting part is that a documented, reproducible SFT+DPO+RLVR pipeline got this far without it. That's the recurring theme in how researchers talk about Tülu: less "best model," more "best paper trail." The training data, preference pairs and the open-instruct codebase are all public, and Ai2's own eval harness ships alongside it, so results are unusually easy to reproduce compared with most frontier post-training work.

The practical complaint is that Tülu was never meant to be used the way a product is used. There's no Ai2-hosted inference endpoint or chat UI; you pull weights from Hugging Face and run them yourself, and the 405B checkpoint needs serious multi-GPU infrastructure that puts it out of reach for anyone without a lab-scale budget. The Llama 3.1-based checkpoints also inherit Meta's Llama Community License restrictions, while the OLMo 2 variant is the fully permissive Apache 2.0 option — a distinction that trips people up if they assume "open" means one license across the board.

Most recently, Ai2 has moved on: OLMo 3 (announced November 2025) absorbed the Tülu post-training approach rather than getting a Tülu 4 label, which means this entry documents a finished, citable milestone rather than a model line still shipping new checkpoints.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Tülu by AllenAI's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Tülu by AllenAI unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Tülu by AllenAI directly →

Spotted something out of date? Suggest an update →

Advertisement