Novita AI

200+ models and per-second GPU cloud, priced low.

LLMs & Chat Paid Has API
Researched · Published · Reviewed
RECATOOLS Score
7.1 / 10
Capability
7
Value for money
8
Ease of use
7
ASEAN readiness
6
API quality
7
Founded
HQ
Users
Launched
Developer

Overview

A unified cloud for serverless LLM/image/audio APIs, dedicated endpoints, and on-demand GPU instances (H100 from $1.99/hr). Aimed at cost-sensitive builders and indie teams.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 13 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

GPU Cloud
From $1.70/GPU-hr
On-demand and spot GPU instances, per-second billing.
  • H100 80GB from $1.99/hr
  • 8-GPU H100 SXM at $1.70/GPU-hr
  • Spot up to 50% off
  • Scale to zero
Dedicated Image
$559/mo
Standard subscription image endpoint.
  • Exclusive high-performance GPU
  • Unlimited images
  • Load 500 models
  • 24/7 support
Dedicated Image Pro
$1,199/mo
Higher-capacity image endpoint.
  • Exclusive high-performance GPU
  • Unlimited images
  • Load 500 models
  • 24/7 support

What you can produce with Novita AI

  • 200+ models across LLM, image, video, audio and embeddings
  • OpenAI-compatible serverless API
  • On-demand GPU instances, H100 from $1.99/GPU-hour
  • Spot instances up to 50% cheaper
  • Serverless GPU that scales to zero, billed per execution
  • Dedicated inference and image endpoints
  • Managed agent sandbox runtimes
Advertisement

ASEAN Perspective

Novita AI in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Two products live under one roof: a serverless API spanning 200+ open models and a rentable GPU cloud billed by the second. Both compete on price. LLM tokens start near $0.02/M, and an H100 runs $1.99/GPU-hour on-demand — with spot instances up to 50% cheaper. The API is OpenAI-compatible and the catalogue is genuinely broad, covering LLM, image, video, audio and embeddings, with new open-weight models added quickly. Where it wobbles is the enterprise-grade stuff: reviewers rate uptime around 3.8/5, serverless SLAs are looser than the dedicated tiers, and prepaid-balance and GPU-limit handling confuses some users. Support runs through Discord rather than formal channels. Treat it as a strong price-optimised supplement or a primary for tolerant workloads — prototyping, batch generation, indie products — rather than the backbone of a latency- and SLA-critical system.

Independent AI-assisted assessment by RECATOOLS.

What people say

Novita's reputation is built on two things developers repeat: low per-token pricing and one API onto a wide model catalogue. Sources put the library at 200+ models across LLM, image, video, audio and embeddings, and Novita scores 4.5/5 on model coverage in one structured review. Cited customers include beBee, Fish Audio, Gizmo and Wiz.ai.

Pricing is aggressive at the small end — Llama 3.1 8B at $0.02/M, Qwen3 4B at $0.03/M — with no monthly minimum and a free start. The GPU side is priced to move: a dedicated H100 80GB endpoint at $1.99/GPU-hour, 8-GPU bare-metal H100 SXM at $1.70/GPU-hr, and spot instances advertising up to 50% savings. Dedicated image endpoints break the usage-based mold as monthly subscriptions — Standard at $559/month, Pro at $1,199/month, both sales-gated.

On experience, reviewers praise fast integration (OpenAI-compatible, usually just a base-URL change), useful docs, and responsive Discord support. The friction shows up around billing mechanics: several users report confusion over prepaid balance and GPU limits.

Reliability is the softer spot. Novita rates 3.8/5 on uptime in one assessment, with dedicated-endpoint SLAs documented at 98%-99.5% availability but serverless guarantees described as less contractual. One vendor comparison cites 99.9% uptime and an order-of-magnitude drop in request-level failures versus weaker providers, so real-world numbers may beat the cautious ratings — but historical incident transparency for procurement is limited. Trustpilot and Product Hunt samples are small and mixed, which makes enterprise satisfaction hard to benchmark.

The net read across sources: Novita is a value play with unusually broad coverage and per-second GPU economics. It rates well on catalogue and price, mid-pack on scaling (4.0/5 performance) and uptime, and light on the enterprise assurances — SLAs, support formality, incident history — that mission-critical buyers want. Serverless GPU that scales to zero and bills only for execution is a genuine draw for spiky workloads.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Novita AI's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Novita AI unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Novita AI directly →

Spotted something out of date? Suggest an update →

Advertisement