Orca

Microsoft Research's reasoning-tuned small models — research-only

LLMs & Chat Open Source Has API Open Source
Researched · Published · Reviewed
RECATOOLS Score
6 / 10
Capability
6
Value for money
8
Ease of use
4
ASEAN readiness
5
API quality
4
Founded
2023
HQ
Redmond, Washington
Users
200k+ downloads
Launched
Jun 2023
Developer
Microsoft

Overview

Microsoft Research's small-model line trained on GPT-4 reasoning traces ('explanation tuning'). Only Orca 2 (7B/13B) was ever released — under a research-only license — and the line has been dormant since 2024, with Microsoft's small-model work now shipping as Phi.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Fully free

Use cases

Research into small model reasoning capabilities using explanation-based training Building reasoning-capable applications that require efficient inference costs Academic study of knowledge distillation from large to small models
Advertisement

ASEAN Perspective

Orca in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Orca is a research programme, not a product. The famous original model was never released; the downloadable Orca 2 checkpoints (7B/13B, Llama-2-based) carry Microsoft's research-only license, ruling out commercial deployment, and Orca-Math and Orca-3 shipped as papers plus datasets, not weights. The ideas were genuinely influential — explanation tuning on GPT-4 reasoning traces is now standard synthetic-data practice, and Orca-Math hit 86.81% on GSM8K with 7B parameters. But the line has been dormant since mid-2024, and Microsoft's active small-model work now ships as the MIT-licensed Phi family. Useful to researchers studying reasoning distillation; irrelevant as a 2026 deployment choice — pick Phi-4 or a current open model instead.

Independent AI-assisted assessment by RECATOOLS.

What people say

The model that made Orca famous was never released. Microsoft Research published the June 2023 paper showing a 13B student model matching GPT-3.5 on reasoning benchmarks after training on GPT-4's step-by-step explanations — then kept the weights. What you can actually download is Orca 2 (7B and 13B, November 2023), Llama-2-based checkpoints on Hugging Face under the Microsoft Research License, which is research-only: no products, no commercial use.

The pattern repeated. Orca-Math (February 2024) pushed a Mistral-7B fine-tune to 86.81% on GSM8K, ahead of Llama-2-70B and GPT-3.5 on grade-school math — Microsoft released the 200k-problem training dataset but not the model. The Orca-3/AgentInstruct paper (July 2024) reported 40–54% gains over Mistral-7B-Instruct on AGIEval, GSM8K and BBH; again, paper only. Since then the line has gone quiet, and Microsoft's small-model effort visibly moved to the Phi family, which ships MIT-licensed weights and gets actual iteration.

Orca's real legacy is the technique. Explanation tuning — distilling reasoning traces rather than final answers — is now standard practice in synthetic-data pipelines, and the progressive-learning recipe shaped how much of the field builds instruction data. For researchers studying distillation, Orca 2 remains a clean, well-documented reference point. As something to deploy in 2026 it's a non-starter: a two-generation-old Llama base, a license that forbids production use, and no maintenance. The score of 6 makes sense only as a nod to influence; as a usable tool it would rate far lower.

Summary of public user & expert reviews, compiled by RECATOOLS.

Notable facts

  • Orca was the first model to systematically train on GPT-4's reasoning traces rather than just final answers — a methodology that became widely adopted across the research community.
  • Orca-Math 7B outperformed GPT-4 on the elementary school math benchmark MATH — a remarkable result for a model 100x smaller.
  • The Orca research paper was cited over 500 times within 6 months of publication, making it one of the most impactful small LLM papers of 2023.

Frequently asked questions

Is Orca free?
Yes. Model weights are free on Hugging Face.
What makes Orca different from other fine-tuned models?
Training on GPT-4's reasoning explanations rather than just final answers teaches the model to reason, not just pattern-match.
Is Orca better than Llama for reasoning?
Yes. Orca's reasoning training specifically improves multi-step logical and mathematical reasoning.
What licence does Orca use?
Microsoft Research Licence — primarily for research use.
What is Orca-Math?
A specialised version of Orca fine-tuned specifically on mathematical problem-solving with GPT-4 reasoning traces.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Orca's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Orca unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Orca directly →

Spotted something out of date? Suggest an update →

Advertisement