Hyperbolic
GPU marketplace and pay-per-token inference for 25+ open AI models
Overview
Hyperbolic runs a decentralized GPU marketplace (RTX 4090 through H200/B200) alongside an OpenAI-compatible serverless inference API for open-weight models like Llama, Qwen and DeepSeek. It's aimed at developers who want hyperscaler-grade compute without hyperscaler pricing or contracts.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- H100/H200/B200/RTX 4090 access
- Serverless inference API billed per token
- No GPU quota limits
- Terms up to 30 days+
- Guaranteed availability
- Volume pricing
- Custom SLAs
- Long-term commitment
What you can produce with Hyperbolic
- OpenAI-compatible serverless inference API for 25+ open models
- On-demand GPU rentals billed by the hour, no commitment
- Reserved GPU clusters at prepaid discount rates
- Private cloud / dedicated long-term GPU infrastructure
- Image generation via Flux and Stable Diffusion, audio via Melo TTS
- Public CLI, AgentKit template and Gradio SDK on GitHub
- Team/organization account management
ASEAN Perspective
Hyperbolic in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Hyperbolic's pitch is simple: aggregate underused GPUs from data centers worldwide and rent them out well below AWS or GCP list price, then wrap that supply in a serverless inference API so developers never touch the hardware layer. The GPU marketplace spans RTX 4090s up to H200/B200 clusters, with reserved and private-cloud tiers for teams that need guaranteed capacity rather than spot pricing.
The catch with any GPU brokerage is durability of supply and uptime under load — Hyperbolic doesn't publish SLAs on par with the hyperscalers, so teams running production traffic should pilot before committing budget. It's best suited to researchers, indie developers and startups optimizing for cost per token or per GPU-hour, not enterprises that need contractual guarantees. The open-source CLI and agent-kit repos suggest a genuinely developer-first shop rather than a marketing wrapper around someone else's cloud.
What people say
Hyperbolic started as a Web3-flavored idea — a decentralized network pooling idle GPU capacity — and has since raised $20M total, including a $12M Series A in December 2024 led by Variant Fund and Polychain Capital, with angels like former Coinbase CTO Balaji Srinivasan and Near's Illia Polosukhin. The company says it now serves over 250,000 engineers and counts Quora, Hugging Face, OpenRouter, Black Forest Labs, Nous Research and LMSYS among its users, plus research groups at Cornell, UC Berkeley, NYU and Stanford.
The product line has three shapes: on-demand GPUs billed hourly with no commitment, reserved clusters at a prepaid discount for sub-year terms, and private cloud for teams that want dedicated long-term infrastructure. Layered on top is a serverless, OpenAI-compatible inference API covering 25+ open models — Llama, Qwen, DeepSeek, Flux for image generation, Melo TTS for audio — so teams that don't want to manage raw GPUs can just call an endpoint. Recent releases include Forge (June 2026, described as the provisioning layer behind GPU reliability), 30-day GPU reservations (February 2026), and dedicated model hosting (January 2026).
Third-party GPU-pricing trackers put Hyperbolic's on-demand rates well under hyperscaler list price — figures cited range roughly $0.35-$3.20 per GPU-hour depending on card and snapshot date, since marketplace pricing moves with supply. That volatility is the trade-off for the discount: rates aren't fixed the way a traditional cloud SKU is, so cost planning needs to account for it rather than assume a static number.
Hyperbolic sits in a genuinely crowded field alongside Together AI, Fireworks and DeepInfra, all chasing the same cost-conscious inference buyer. What differentiates it somewhat is the open developer tooling — a public CLI, an AgentKit template repo, and a Gradio integration package are all on GitHub under HyperbolicLabs, which is more transparency than most GPU brokers offer. No independent Reddit or review-site consensus turned up in research; the available signal is mostly the company's own funding and customer-logo announcements plus GPU-pricing aggregator data.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Hyperbolic's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Hyperbolic unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Hyperbolic directly →
Spotted something out of date? Suggest an update →
Alternatives to Hyperbolic
More in LLMs & Chat