Groq
Ultra-fast open-weight model inference on custom LPU silicon
Overview
Groq is an AI inference company built around its Language Processing Unit (LPU), a chip architecture designed for extremely low-latency, high-throughput LLM serving. After licensing much of its chip IP to Nvidia in late 2025, Groq raised $650M in June 2026 to bet fully on its GroqCloud inference platform, targeting roughly 200 MW of capacity by 2027.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 24 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 14,400 requests/day, 30 RPM, 6,000 TPM
- All open-weight models included
- Limits apply at the org level
- Llama 3.1 8B: $0.05 in / $0.08 out per 1M
- Llama 3.3 70B: $0.59 in / $0.79 out per 1M
- Prompt caching: 50% off cached input
- Batch API: 50% lower cost
- Reserved capacity and higher limits
- Custom terms via sales
Use cases
What you can produce with Groq
- Build real-time voice AI pipelines combining Whisper Large v3 Turbo transcription (~216x real-time) with Llama generation for low end-to-end latency
- Prototype LLM features instantly on the free tier using a simple base-URL swap in any existing OpenAI SDK integration
- Deploy streaming chat interfaces where Groq's market-leading throughput (~276 t/s on Llama 3.3 70B, per Artificial Analysis) minimises perceptible generation lag
- Run autonomous AI agents via Compound Beta, which bundles built-in web search and code execution with no external tool wiring
- Cut inference costs on large-scale summarisation, translation, or classification jobs using the Batch API for async processing
- Build multilingual ASEAN content tools on Llama 3.3 70B or Qwen 2.5 with lower latency via the Sydney APAC inference node
- Implement content moderation pipelines using Llama Guard models served at Groq throughput rates for near-real-time safety checks
ASEAN Perspective
Groq in Southeast Asia
Groq is gaining meaningful traction in Southeast Asia, driven by a price-to-speed ratio that appeals to cost-conscious startups in Singapore, Indonesia, Malaysia, and the Philippines. Groq says over half of its developers are Asia-Pacific based, and the Sydney APAC inference node (opened November 2025 with Equinix) reduces latency for the region versus US-only endpoints. For ASEAN developers building voice AI, multilingual chatbots, or real-time agent workflows on open-weight models like Llama or Qwen, Groq's free tier and low cost per million tokens make it one of the most accessible inference options available. The absence of Singapore or Southeast Asian data centres means compliance-sensitive workloads (finance, healthcare) still carry data-residency concerns, though Groq has signalled Asia as an active expansion target. Groq does not serve proprietary models, limiting its fit for ASEAN enterprises standardised on OpenAI or Anthropic. A further caveat: after Nvidia's December 2025 deal absorbed Groq's hardware assets and founding team, the independent GroqCloud is rebuilding as an inference neocloud — its API is operating normally today, but its long-term roadmap is worth monitoring. For pure-speed inference on open-weight models, it remains hard to beat at this price point.
Groq is among the fastest options for open-weight model inference — if raw tokens-per-second throughput is your primary requirement, few providers beat it at this price point (independent benchmarks put Llama 3.3 70B at ~276 t/s, the fastest measured). The OpenAI-compatible API, generous free tier, and predictable per-token pricing make it genuinely easy to adopt. However, the hard ceiling of open-weight-only models means teams tied to GPT-4o, Claude, or Gemini cannot use Groq at all, and the free tier's 6,000 TPM cap forces paid upgrades sooner than it appears. The bigger caveat is strategic: Nvidia's December 2025 deal absorbed Groq's chip assets and most of its senior team (including founder Jonathan Ross), leaving the independent GroqCloud to rebuild as an inference 'neocloud' under new leadership — its long-term hardware roadmap deserves scrutiny even as the API operates normally today.
What people say
Ask developers what Groq is for and you get one answer: speed. Llama 3.3 70B at roughly 276 tokens per second — the fastest Artificial Analysis has measured for that model — feels instantaneous next to GPU-backed rivals, and the OpenAI-compatible API means migrating is usually a base-URL change. The free tier needs no credit card, and GroqCloud claimed more than 3.5 million developers by February 2026.
The complaints are just as consistent. The catalogue is open-weight only — no GPT, Claude or Gemini — a hard blocker for teams tied to proprietary models. Free-tier ceilings (30 requests and 6,000 tokens per minute on most models, plus daily caps) bite fast in production; one long prompt can eat a minute's budget, and limits apply org-wide. The paid developer tier lifts limits roughly 10x for the price of adding a card. And Groq trains nothing itself, so quality tracks whatever Meta, Google, Mistral and Alibaba release as open weights.
The strategic question is newer. In December 2025 Nvidia paid about $20 billion to license Groq's LPU technology and hire most of its senior team, founder Jonathan Ross included. GroqCloud was carved out and stays independent under former CFO Simon Edwards, now rebuilding as an inference 'neocloud' on remaining LPU hardware plus Nvidia GPUs, funded by a roughly $650 million raise reported in late May 2026. The API runs uninterrupted today, but the long-term hardware roadmap deserves scrutiny.
User verdict: a superb specialist for voice AI, streaming agents and rapid prototyping — not a one-stop inference platform.
Summary of public user & expert reviews, compiled by RECATOOLS.
Notable facts
- Groq founder Jonathan Ross originally created the Google TPU as a '20% side project' (one day a week) before leaving Google in 2016 to build LPU chips — which Nvidia eventually acquired/licensed in a ~$20 billion deal in December 2025, taking Ross and most of Groq's senior team with it
- Groq's LPU uses on-chip SRAM instead of off-chip DRAM for model weights, enabling deterministic synchronous dataflow that eliminates memory-bandwidth stalls — the architectural reason it tops independent Artificial Analysis throughput benchmarks (~276 tokens/second on Llama 3.3 70B, the fastest of all measured providers)
Frequently asked questions
About this listing
This entry was compiled from publicly available data including Groq's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Groq unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Groq directly →
Spotted something out of date? Suggest an update →
More in Code & Dev Tools