Airavata
India's first open instruction-tuned Hindi LLM, built by IIT Madras's AI4Bharat to advance Indic language AI for 1.4 billion people.
Overview
Airavata is a 7-billion-parameter instruction-tuned language model for Hindi, released in January 2024 by AI4Bharat — the AI research lab incubated at IIT Madras and backed by Nandan Nilekani's philanthropic grants (totalling INR 70 crore) and venture funding from Peak XV and Lightspeed. It is built by fine-tuning Sarvam AI's OpenHathi (itself derived from Meta's LLaMA 2) on the IndicInstruct dataset — 385,000 curated instruction-response pairs spanning Flan-v2, Dolly, OpenAssistant, Anthropic-HHH, wikiHow, and the team's own Anudesh corpus — without using any data generated by proprietary models such as GPT-4, enabling license-friendly downstream use.
The model is fully open-weights (under the LLaMA 2 license), available on Hugging Face with the IndicInstruct training code on GitHub. It outperforms the base OpenHathi model on Hindi NLU tasks (IndicSentiment F1: 97.01; IndicXNLI F1: 74.7) and produces more naturally fluent Hindi than GPT-4 in human evaluations, while scoring 43.90 on MMLU — comparable to similarly sized 7B English LLMs. AI4Bharat has publicly committed to expanding coverage to all 22 constitutionally scheduled Indian languages; as of mid-2026, Airavata remains the flagship Hindi chat model in the ecosystem, with its models embedded in India's national Bhashini language platform under the IndiaAI Mission.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- Llama 2 licence (via OpenHathi base)
- IndicInstruct fine-tune, research-focused
- No sign-up cost or usage fees
- Works with vLLM, Ollama, LM Studio
- No rate limits or quotas
- No hosted API or vendor support
Use cases
What you can produce with Airavata
- Hindi instruction-following chat completions via local inference (vLLM, Ollama, LM Studio)
- Open IndicInstruct dataset of curated Hindi and English instruction pairs for fine-tuning
- Reproducible LoRA fine-tuning pipeline via the AI4Bharat/IndicInstruct GitHub repository
- Airavata Evaluation Suite on HuggingFace for benchmarking Hindi LLM capabilities
- Integration reference for Bhashini / IndiaAI Mission sovereign language platform
- IndicSentiment and IndicXNLI benchmark baselines for Hindi NLU tasks
ASEAN Perspective
Airavata in Southeast Asia
Airavata is an India-first model with limited direct ASEAN language coverage, as it targets Hindi and other scheduled Indian languages rather than Malay, Thai, Vietnamese, Tagalog, or Bahasa Indonesia. However, it is academically significant for ASEAN NLP researchers: its IndicInstruct methodology — clean, GPT-4-data-free instruction tuning on a 7B base — is a directly transferable blueprint for building instruction-tuned models in low-resource Southeast Asian languages. Singapore's A*STAR was a collaborating institution on the Airavata paper, indicating regional research ties. Teams at AI4Bharat and Singapore's research ecosystem have co-published, making Airavata a useful reference point even for APAC language technology builders who are not targeting Hindi.
Airavata is a landmark academic contribution that demonstrates how a well-curated, license-clean instruction dataset can lift a 7B Hindi model above much larger commercial models on naturalness — a result that matters enormously for India's 500 million+ Hindi speakers. Its open-weights release under a licence-friendly framework, combined with the IndicInstruct dataset and evaluation suite, gives researchers and developers a reproducible foundation to build on, and its integration into India's national Bhashini platform shows real sovereign AI deployment at scale.
The practical caveats are significant for production use: there is no managed API, no safety alignment layer beyond basic dataset curation, and the LLaMA 2 licence (rather than Apache 2.0) introduces commercial-use friction. Coverage remains Hindi-primary as of mid-2026 — the promised 22-language expansion is a research roadmap, not a delivered product. ASEAN developers will find limited direct utility unless they are specifically working on Hindi-language features or want a reference implementation of Indic instruction tuning to adapt for their own low-resource language.
What people say
Seven billion parameters and 385,000 hand-curated instruction pairs — that's the entire recipe behind Airavata, and it's enough to beat GPT-4 on fluency in blind Hindi evaluations. AI4Bharat, the IIT Madras lab that built it, fine-tuned Sarvam AI's OpenHathi (itself a LLaMA 2 derivative) on IndicInstruct, a dataset assembled from Flan-v2, Dolly, OpenAssistant, Anthropic-HHH and the team's own Anudesh corpus, deliberately avoiding any GPT-4-generated training data to keep the license clean.
The scores hold up: 97.01 F1 on IndicSentiment, 74.7 on IndicXNLI, and an MMLU of 43.90 that's respectable for a 7B model but nowhere near frontier territory. What Airavata proves isn't raw power — it's that a small, well-curated instruction set beats brute-force scale for Hindi naturalness, a finding that's since been cited across a wave of Indic-language papers and folded into India's national Bhashini platform under the IndiaAI Mission.
None of that makes it a product. There's no hosted API, no chat interface, no safety layer beyond what the base data provides, and the LLaMA 2 license (not Apache) makes commercial use a legal headache rather than a checkbox. AI4Bharat's promise to cover all 22 scheduled Indian languages is still a roadmap item two years on — Hindi remains the only language actually shipped. Weights and training code are free on Hugging Face and GitHub, funded in part by Nandan Nilekani's philanthropic grants (INR 70 crore and counting) to the Nilekani Centre at AI4Bharat. Treat Airavata as a research baseline to fork, not a service to integrate — researchers and government AI teams get the most out of it; anyone wanting a Hindi chatbot out of the box should look elsewhere.
Summary of public user & expert reviews, compiled by RECATOOLS.
Notable facts
- Airavata is named after the mythical white elephant of Hindu lore — a fitting choice for a model designed to carry the weight of Hindi AI into the future.
- The IndicInstruct dataset contains 74.7 million prompt-response pairs across 20 Indian languages, one of the largest open instruction corpora ever built for non-English languages.
- Singapore's A*STAR was a co-author on the Airavata paper — a quiet sign of ASEAN's academic investment in South Asian language AI.
- Nandan Nilekani, the architect of Aadhaar (India's 1.3-billion-person biometric ID system), has committed INR 70 crore in philanthropic grants to AI4Bharat, the lab behind Airavata.
Frequently asked questions
About this listing
This entry was compiled from publicly available data including Airavata's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Airavata unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Airavata directly →
Spotted something out of date? Suggest an update →
More in LLMs & Chat