RunPod
Per-second GPU cloud — Pods, serverless inference, and multi-node clusters
Overview
On-demand and spot GPU rental for AI workloads, billed by the second. Dedicated Pods for training, Serverless endpoints that scale to zero when idle, and Instant Clusters for distributed jobs — popular for self-hosting open-weight LLMs and diffusion models without a long-term cloud contract.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 13 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
What you can produce with RunPod
- On-demand and spot GPU Pods (RTX 4090, A100, H100) billed per second
- Serverless auto-scaling inference endpoints that scale to zero when idle
- Instant Clusters for multi-node distributed training
- Secure Cloud (datacenter) and Community Cloud (~50% cheaper) options
- Pre-built templates for common ML, LLM and diffusion stacks
- CLI (runpodctl), Python SDK and MCP server
ASEAN Perspective
RunPod in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
The pitch that keeps RunPod near the top of GPU-cloud shortlists is honest billing: you pay by the second, and Serverless endpoints cost nothing while idle, so a model can become a product without paying for a GPU that sits waiting. Pre-built templates remove most of the CUDA-and-drivers setup that scares newcomers off. Two things to weigh. Community Cloud rents consumer GPUs from third parties at roughly half the Secure Cloud rate, but those instances can be preempted without warning — fine for experiments, risky for anything customer-facing (measured 18-month uptime ran ~99.71% Secure vs ~97.98% Community). And Serverless cold starts of 30+ seconds hurt latency-sensitive apps unless you keep workers warm. Best fit for indie ML engineers, researchers and startups who want cheap, flexible GPUs; less compelling if you're already deep in a hyperscaler's ecosystem or need enterprise SLAs.
What people say
Reviewers in 2026 tend to crown RunPod the best all-round GPU cloud for small teams, and the reasons are consistent: per-second billing, low headline rates, and templates that get a model running fast.
On price, published Secure Cloud rates run from about $0.69/hr for an RTX 4090, $1.49/hr for an A100 and $2.89/hr for an H100, with Community Cloud roughly 50% cheaper — the Community H100 SXM at ~$2.39/hr is frequently called the cheapest published rate for that SKU. Overall the platform spans roughly $0.27 to $7.39 per GPU-hour depending on tier and card. There's no permanent free tier, though signup bonus credits and a no-VC-required startup program (up to ~$1,000 credit) lower the barrier.
The headline caveat is reliability by tier. Community Cloud uses consumer GPUs from third-party hosts and instances can be preempted without notice, which reviewers flag as unsuitable for production. A blended 18-month uptime figure cited across comparisons puts Secure Cloud around 99.71% and Community around 97.98% — a real gap when uptime matters.
Serverless draws the most nuanced feedback. It's described as up to ~87% cheaper when workloads are genuinely bursty, because you pay only while requests run. But cold starts of 30+ seconds come up repeatedly, and reviewers note it can end up pricier than an always-on Pod if traffic is steady rather than spiky. The practical advice is to match the product to the pattern: Pods for sustained load, Serverless for spiky inference, Clusters for distributed training.
The developer experience gets steady praise — templates, a usable CLI (runpodctl), a Python SDK and an MCP server mean you can go from zero to a served model quickly. Criticism clusters around support depth and the occasional capacity crunch on popular GPUs.
Net read from the review corpus: strong value and flexibility for independent engineers and startups self-hosting open-weight models, with the important asterisk that Community Cloud's savings come with preemption risk, and Serverless economics only work when the workload is truly bursty.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including RunPod's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with RunPod unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to RunPod directly →
Spotted something out of date? Suggest an update →
More in Code & Dev Tools