Cohere
Enterprise and sovereign AI — RAG-grade models, on-prem deployment
Overview
Enterprise AI company building the Command, Embed and Rerank model families for RAG, search and agents — deployable in its cloud, on the hyperscalers, or fully on-prem. Merged with Germany's Aleph Alpha in April 2026 in a sovereign-AI push valued around $20B.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 1,000 free API calls/mo across endpoints
- Non-commercial use only
- Chat limited to 20 req/min
- Command A: $2.50/M input, $10/M output
- Command R: $0.50/M input, $1.50/M output
- Embed and Rerank models available
- Billed monthly or at $250 balance
- Model Vault dedicated instances (from $2,500/mo)
- North workplace AI platform
- Private and on-prem deployment options
Use cases
What you can produce with Cohere
- Build a production RAG pipeline using Cohere Embed v4 to convert documents into semantic vectors and Rerank 4 to re-score retrieved chunks before passing them to a language model, improving answer accuracy in enterprise search.
- Generate structured summaries, classifications, and entity extractions from large document batches using the Command A API (111B dense model, 256K-token context window) — suitable for legal, financial, and compliance document workflows.
- Fine-tune a custom Command model on proprietary company data via Cohere's fine-tuning API to adapt tone, terminology, and output format to a specific enterprise domain without exposing data to external hyperscalers.
- Deploy Cohere models on sovereign or private cloud infrastructure (AWS, Azure, GCP, or on-premises) to meet data-residency requirements for government, healthcare, or defence use cases — a core differentiator from OpenAI and Anthropic.
- Prototype and test multilingual NLP features using the Trial API key (1,000 free calls/month, non-commercial, 20 req/min Chat limit) before committing to a paid production key.
- Run multi-step agentic workflows — including tool-use orchestration and complex PDF and chart analysis — using Command A+ (218B sparse MoE, 25B active parameters, 128K context, runs on as few as two H100 GPUs at W4A4 quantisation) released May 2026 under Apache 2.0.
- Embed semantic search into a SaaS product or internal tool using Cohere's managed Embed endpoint (Embed v4 at $0.12/1M text tokens), paying per token rather than running self-hosted vector infrastructure.
ASEAN Perspective
Cohere in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Cohere's pitch is deployment control, not frontier benchmarks. Embed and Rerank remain among the strongest retrieval models for RAG and enterprise search, and the Command line was refreshed in May 2026 with Command A+, a 218B mixture-of-experts model released under Apache 2.0. Models run in Cohere's cloud, on AWS, Azure or GCP, or fully on-prem — the reason regulated and public-sector buyers pick it — and the April 2026 merger with Aleph Alpha, at a combined valuation near $20B, doubles down on sovereign AI. There's no consumer chatbot, pure reasoning trails OpenAI, Anthropic and Google, and per-token prices look steep beside open-weight rivals. For ASEAN enterprises needing private deployment and multilingual embeddings it's a strong, often-overlooked pick; casual users should look elsewhere.
What people say
$20 billion is the number that redefined Cohere this year. The April 2026 merger with Germany's Aleph Alpha — anchored by roughly $600M from retail giant Schwarz Group and announced with both the Canadian and German digital ministers in the room — turned a respected also-ran into the largest sovereign-AI vendor outside Silicon Valley. The bet isn't beating OpenAI on frontier reasoning; it's being the model provider that governments and regulated industries can deploy inside their own walls.
The technical reputation rests on retrieval. Practitioners consistently rank Embed and Rerank among the best components you can put in a RAG pipeline, and adding Cohere's reranker on top of plain vector search is one of the few upgrades that reliably shows up in accuracy numbers. The Command line moved fast in 2026: Command R/R+ was retired in April, and Command A+ arrived on May 20 — a 218B-total, 25B-active mixture-of-experts model, notable as the family's first Apache 2.0-licensed frontier release, aimed squarely at agentic tool use and multilingual enterprise work.
The criticisms haven't changed. No consumer chatbot means thin developer mindshare next to OpenAI and Anthropic. Command A's $2.50 input / $10 output per-million-token pricing is fine inside enterprise contracts but looks expensive to small teams benchmarking against Mistral or open-weight alternatives. And the Aleph Alpha integration is a promise, not a product — the deal was announced in April, and how the two stacks merge remains unspecified.
For ASEAN enterprises with data-residency requirements, though, the VPC and on-prem deployment options plus genuinely strong multilingual embeddings make Cohere one of the few credible non-hyperscaler choices. Casual users have no reason to be here, and Cohere seems fine with that.
Summary of public user & expert reviews, compiled by RECATOOLS.
Notable facts
- Cohere co-founder Aidan Gomez was a co-author of the 'Attention Is All You Need' paper that introduced the Transformer architecture, which underpins every modern LLM.
- Cohere's models can be deployed entirely within a customer's own AWS or Azure environment, with data never leaving the enterprise boundary.
- The company's Command R+ achieved first place on the Berkeley Function-Calling Leaderboard, outperforming GPT-4 on structured tool use tasks.
Frequently asked questions
About this listing
This entry was compiled from publicly available data including Cohere's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Cohere unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Cohere directly →
Spotted something out of date? Suggest an update →
More in LLMs & Chat