Airavata

India's first open instruction-tuned Hindi LLM, built by IIT Madras's AI4Bharat to advance Indic language AI for 1.4 billion people.

LLMs & Chat Open Source Open Source
Researched · Published · Reviewed
RECATOOLS Score
6.8 / 10
Capability
6.5
Value for money
9.5
Ease of use
5
ASEAN readiness
3.5
API quality
2
Founded
2020
HQ
Chennai, India (IIT Madras)
Users
Research and developer community; open-weights model on HuggingFace
Launched
January 2024 (arXiv 2401.15006 + HuggingFace)
Developer
AI4Bharat / IIT Madras (Nilekani Centre at AI4Bharat)

Overview

Airavata is a 7-billion-parameter instruction-tuned language model for Hindi, released in January 2024 by AI4Bharat — the AI research lab incubated at IIT Madras and backed by Nandan Nilekani's philanthropic grants (totalling INR 70 crore) and venture funding from Peak XV and Lightspeed. It is built by fine-tuning Sarvam AI's OpenHathi (itself derived from Meta's LLaMA 2) on the IndicInstruct dataset — 385,000 curated instruction-response pairs spanning Flan-v2, Dolly, OpenAssistant, Anthropic-HHH, wikiHow, and the team's own Anudesh corpus — without using any data generated by proprietary models such as GPT-4, enabling license-friendly downstream use.

The model is fully open-weights (under the LLaMA 2 license), available on Hugging Face with the IndicInstruct training code on GitHub. It outperforms the base OpenHathi model on Hindi NLU tasks (IndicSentiment F1: 97.01; IndicXNLI F1: 74.7) and produces more naturally fluent Hindi than GPT-4 in human evaluations, while scoring 43.90 on MMLU — comparable to similarly sized 7B English LLMs. AI4Bharat has publicly committed to expanding coverage to all 22 constitutionally scheduled Indian languages; as of mid-2026, Airavata remains the flagship Hindi chat model in the ecosystem, with its models embedded in India's national Bhashini language platform under the IndiaAI Mission.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Self-hosted
Free
Run locally on your own hardware — compute costs are yours
  • Works with vLLM, Ollama, LM Studio
  • No rate limits or quotas
  • No hosted API or vendor support

Use cases

Building Hindi-language chatbots and virtual assistants for Indian government services Fine-tuning a Hindi instruction base for domain-specific applications (healthcare, finance, agriculture) Academic research into low-resource and Indic language NLP Evaluating and benchmarking other multilingual LLMs on Hindi tasks using the IndicInstruct evaluation suite Adapting the IndicInstruct methodology to build instruction-tuned models for other low-resource languages in South or Southeast Asia

What you can produce with Airavata

  • Hindi instruction-following chat completions via local inference (vLLM, Ollama, LM Studio)
  • Open IndicInstruct dataset of curated Hindi and English instruction pairs for fine-tuning
  • Reproducible LoRA fine-tuning pipeline via the AI4Bharat/IndicInstruct GitHub repository
  • Airavata Evaluation Suite on HuggingFace for benchmarking Hindi LLM capabilities
  • Integration reference for Bhashini / IndiaAI Mission sovereign language platform
  • IndicSentiment and IndicXNLI benchmark baselines for Hindi NLU tasks
Advertisement

ASEAN Perspective

Airavata in Southeast Asia

Airavata is an India-first model with limited direct ASEAN language coverage, as it targets Hindi and other scheduled Indian languages rather than Malay, Thai, Vietnamese, Tagalog, or Bahasa Indonesia. However, it is academically significant for ASEAN NLP researchers: its IndicInstruct methodology — clean, GPT-4-data-free instruction tuning on a 7B base — is a directly transferable blueprint for building instruction-tuned models in low-resource Southeast Asian languages. Singapore's A*STAR was a collaborating institution on the Airavata paper, indicating regional research ties. Teams at AI4Bharat and Singapore's research ecosystem have co-published, making Airavata a useful reference point even for APAC language technology builders who are not targeting Hindi.

RECATOOLS Verdict

Airavata is a landmark academic contribution that demonstrates how a well-curated, license-clean instruction dataset can lift a 7B Hindi model above much larger commercial models on naturalness — a result that matters enormously for India's 500 million+ Hindi speakers. Its open-weights release under a licence-friendly framework, combined with the IndicInstruct dataset and evaluation suite, gives researchers and developers a reproducible foundation to build on, and its integration into India's national Bhashini platform shows real sovereign AI deployment at scale.

The practical caveats are significant for production use: there is no managed API, no safety alignment layer beyond basic dataset curation, and the LLaMA 2 licence (rather than Apache 2.0) introduces commercial-use friction. Coverage remains Hindi-primary as of mid-2026 — the promised 22-language expansion is a research roadmap, not a delivered product. ASEAN developers will find limited direct utility unless they are specifically working on Hindi-language features or want a reference implementation of Indic instruction tuning to adapt for their own low-resource language.

Independent AI-assisted assessment by RECATOOLS.

What people say

Seven billion parameters and 385,000 hand-curated instruction pairs — that's the entire recipe behind Airavata, and it's enough to beat GPT-4 on fluency in blind Hindi evaluations. AI4Bharat, the IIT Madras lab that built it, fine-tuned Sarvam AI's OpenHathi (itself a LLaMA 2 derivative) on IndicInstruct, a dataset assembled from Flan-v2, Dolly, OpenAssistant, Anthropic-HHH and the team's own Anudesh corpus, deliberately avoiding any GPT-4-generated training data to keep the license clean.

The scores hold up: 97.01 F1 on IndicSentiment, 74.7 on IndicXNLI, and an MMLU of 43.90 that's respectable for a 7B model but nowhere near frontier territory. What Airavata proves isn't raw power — it's that a small, well-curated instruction set beats brute-force scale for Hindi naturalness, a finding that's since been cited across a wave of Indic-language papers and folded into India's national Bhashini platform under the IndiaAI Mission.

None of that makes it a product. There's no hosted API, no chat interface, no safety layer beyond what the base data provides, and the LLaMA 2 license (not Apache) makes commercial use a legal headache rather than a checkbox. AI4Bharat's promise to cover all 22 scheduled Indian languages is still a roadmap item two years on — Hindi remains the only language actually shipped. Weights and training code are free on Hugging Face and GitHub, funded in part by Nandan Nilekani's philanthropic grants (INR 70 crore and counting) to the Nilekani Centre at AI4Bharat. Treat Airavata as a research baseline to fork, not a service to integrate — researchers and government AI teams get the most out of it; anyone wanting a Hindi chatbot out of the box should look elsewhere.

Summary of public user & expert reviews, compiled by RECATOOLS.

Notable facts

  • Airavata is named after the mythical white elephant of Hindu lore — a fitting choice for a model designed to carry the weight of Hindi AI into the future.
  • The IndicInstruct dataset contains 74.7 million prompt-response pairs across 20 Indian languages, one of the largest open instruction corpora ever built for non-English languages.
  • Singapore's A*STAR was a co-author on the Airavata paper — a quiet sign of ASEAN's academic investment in South Asian language AI.
  • Nandan Nilekani, the architect of Aadhaar (India's 1.3-billion-person biometric ID system), has committed INR 70 crore in philanthropic grants to AI4Bharat, the lab behind Airavata.

Frequently asked questions

Is Airavata free to use commercially?
Airavata's weights are open but governed by the LLaMA 2 licence, which restricts use in products serving more than 700 million monthly users and requires a separate Meta commercial licence for large-scale deployments. For most startups and researchers this is not a barrier, but enterprise teams should review Meta's LLaMA 2 licence terms before deploying.
Does Airavata support languages other than Hindi?
The v0.1 release targets Hindi as its primary language, with retained English capability from its LLaMA 2 foundation. AI4Bharat has publicly committed to expanding Airavata to all 22 constitutionally scheduled Indian languages, but as of mid-2026 a multi-language production release has not been confirmed. The broader AI4Bharat ecosystem (IndicTrans2, IndicASR) covers all 22 languages.
Can I access Airavata via an API without self-hosting?
There is no official managed API for Airavata as of mid-2026. The model weights are available on HuggingFace and can be run locally via vLLM, Ollama, or LM Studio. Developers looking for a hosted Indic LLM API should explore Sarvam AI or the Bhashini platform, both of which draw on AI4Bharat's underlying research.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Airavata's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Airavata unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Airavata directly →

Spotted something out of date? Suggest an update →

Advertisement