Novita AI
200+ models and per-second GPU cloud, priced low.
Overview
A unified cloud for serverless LLM/image/audio APIs, dedicated endpoints, and on-demand GPU instances (H100 from $1.99/hr). Aimed at cost-sensitive builders and indie teams.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 13 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 200+ models
- OpenAI-compatible
- No monthly minimum
- Free to start
- H100 80GB from $1.99/hr
- 8-GPU H100 SXM at $1.70/GPU-hr
- Spot up to 50% off
- Scale to zero
- Exclusive high-performance GPU
- Unlimited images
- Load 500 models
- 24/7 support
- Exclusive high-performance GPU
- Unlimited images
- Load 500 models
- 24/7 support
What you can produce with Novita AI
- 200+ models across LLM, image, video, audio and embeddings
- OpenAI-compatible serverless API
- On-demand GPU instances, H100 from $1.99/GPU-hour
- Spot instances up to 50% cheaper
- Serverless GPU that scales to zero, billed per execution
- Dedicated inference and image endpoints
- Managed agent sandbox runtimes
ASEAN Perspective
Novita AI in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Two products live under one roof: a serverless API spanning 200+ open models and a rentable GPU cloud billed by the second. Both compete on price. LLM tokens start near $0.02/M, and an H100 runs $1.99/GPU-hour on-demand — with spot instances up to 50% cheaper. The API is OpenAI-compatible and the catalogue is genuinely broad, covering LLM, image, video, audio and embeddings, with new open-weight models added quickly. Where it wobbles is the enterprise-grade stuff: reviewers rate uptime around 3.8/5, serverless SLAs are looser than the dedicated tiers, and prepaid-balance and GPU-limit handling confuses some users. Support runs through Discord rather than formal channels. Treat it as a strong price-optimised supplement or a primary for tolerant workloads — prototyping, batch generation, indie products — rather than the backbone of a latency- and SLA-critical system.
What people say
Novita's reputation is built on two things developers repeat: low per-token pricing and one API onto a wide model catalogue. Sources put the library at 200+ models across LLM, image, video, audio and embeddings, and Novita scores 4.5/5 on model coverage in one structured review. Cited customers include beBee, Fish Audio, Gizmo and Wiz.ai.
Pricing is aggressive at the small end — Llama 3.1 8B at $0.02/M, Qwen3 4B at $0.03/M — with no monthly minimum and a free start. The GPU side is priced to move: a dedicated H100 80GB endpoint at $1.99/GPU-hour, 8-GPU bare-metal H100 SXM at $1.70/GPU-hr, and spot instances advertising up to 50% savings. Dedicated image endpoints break the usage-based mold as monthly subscriptions — Standard at $559/month, Pro at $1,199/month, both sales-gated.
On experience, reviewers praise fast integration (OpenAI-compatible, usually just a base-URL change), useful docs, and responsive Discord support. The friction shows up around billing mechanics: several users report confusion over prepaid balance and GPU limits.
Reliability is the softer spot. Novita rates 3.8/5 on uptime in one assessment, with dedicated-endpoint SLAs documented at 98%-99.5% availability but serverless guarantees described as less contractual. One vendor comparison cites 99.9% uptime and an order-of-magnitude drop in request-level failures versus weaker providers, so real-world numbers may beat the cautious ratings — but historical incident transparency for procurement is limited. Trustpilot and Product Hunt samples are small and mixed, which makes enterprise satisfaction hard to benchmark.
The net read across sources: Novita is a value play with unusually broad coverage and per-second GPU economics. It rates well on catalogue and price, mid-pack on scaling (4.0/5 performance) and uptime, and light on the enterprise assurances — SLAs, support formality, incident history — that mission-critical buyers want. Serverless GPU that scales to zero and bills only for execution is a genuine draw for spiky workloads.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Novita AI's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Novita AI unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Novita AI directly →
Spotted something out of date? Suggest an update →
Novita AI in the news
Cloud & Infra
Wafr's $100M Cooling Raise Is Small — the AI Data-Centre Water Problem It Targets Is Not
Cloud & Infra
Meta Compute: Meta Is Reportedly Exploring a Cloud Business for Excess AI Capacity
Cloud & Infra
The Cloud Market Just Crossed a Half-Trillion-Dollar Run Rate — and AI Is Why It's Still A...
Alternatives to Novita AI
More in LLMs & Chat