Novita AI
200+ models and per-second GPU cloud, priced low.
Overview
A unified cloud for serverless LLM/image/audio APIs, dedicated endpoints, and on-demand GPU instances (H100 from $1.70/hr). Aimed at cost-sensitive builders and indie teams.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 3 Sep 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 200+ models
- OpenAI-compatible
- No monthly minimum
- Free to start
- H100 80GB from $1.99/hr
- 8-GPU H100 SXM at $1.70/GPU-hr
- Spot up to 50% off
- Scale to zero
- Exclusive high-performance GPU
- Unlimited images
- Load 500 models
- 24/7 support
- Exclusive high-performance GPU
- Unlimited images
- Load 500 models
- 24/7 support
What you can produce with Novita AI
- 200+ models across LLM, image, video, audio and embeddings
- OpenAI-compatible serverless API
- On-demand GPU instances, H100 from $1.99/GPU-hour
- Spot instances up to 50% cheaper
- Serverless GPU that scales to zero, billed per execution
- Dedicated inference and image endpoints
- Managed agent sandbox runtimes
Two products live under one roof: a serverless API spanning 200+ open models and a rentable GPU cloud billed by the second. Both compete on price. LLM tokens start near $0.02/M, and an H100 runs $1.99/GPU-hour on-demand — with spot instances up to 50% cheaper. The API is OpenAI-compatible and the catalogue is genuinely broad, covering LLM, image, video, audio and embeddings, with new open-weight models added quickly. Where it wobbles is the enterprise-grade stuff: reviewers rate uptime around 3.8/5, serverless SLAs are looser than the dedicated tiers, and prepaid-balance and GPU-limit handling confuses some users. Support runs through Discord rather than formal channels. Treat it as a strong price-optimised supplement or a primary for tolerant workloads — prototyping, batch generation, indie products — rather than the backbone of a latency- and SLA-critical system.
What people say
Novita's reputation is built on two things developers repeat: low per-token pricing and one API onto a wide model catalogue. Sources put the library at 200+ models across LLM, image, video, audio and embeddings, and Novita scores 4.5/5 on model coverage in one structured review. Cited customers include beBee, Fish Audio, Gizmo and Wiz.ai.
Pricing is aggressive at the small end — Llama 3.1 8B at $0.02/M, Qwen3 4B at $0.03/M — with no monthly minimum and a free start. The GPU side is priced to move: a dedicated H100 80GB endpoint at $1.99/GPU-hour, 8-GPU bare-metal H100 SXM at $1.70/GPU-hr, and spot instances advertising up to 50% savings. Dedicated image endpoints break the usage-based mold as monthly subscriptions — Standard at $559/month, Pro at $1,199/month, both sales-gated.
On experience, reviewers praise fast integration (OpenAI-compatible, usually just a base-URL change), useful docs, and responsive Discord support. The friction shows up around billing mechanics: several users report confusion over prepaid balance and GPU limits.
Reliability is the softer spot. Novita rates 3.8/5 on uptime in one assessment, with dedicated-endpoint SLAs documented at 98%-99.5% availability but serverless guarantees described as less contractual. One vendor comparison cites 99.9% uptime and an order-of-magnitude drop in request-level failures versus weaker providers, so real-world numbers may beat the cautious ratings — but historical incident transparency for procurement is limited. Trustpilot and Product Hunt samples are small and mixed, which makes enterprise satisfaction hard to benchmark.
The net read across sources: Novita is a value play with unusually broad coverage and per-second GPU economics. It rates well on catalogue and price, mid-pack on scaling (4.0/5 performance) and uptime, and light on the enterprise assurances — SLAs, support formality, incident history — that mission-critical buyers want. Serverless GPU that scales to zero and bills only for execution is a genuine draw for spiky workloads.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Novita AI's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Novita AI unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Novita AI directly →
Spotted something out of date? Suggest an update →
Alternatives to Novita AI
More in LLMs & Chat