Inference
Production model serving that hit $600M ARR in 2026.
1,800+ tokens/sec inference on wafer-scale silicon.
Open-model inference from $0.02 per million tokens.
Flat monthly rate for unlimited use of 20,000+ open LLMs
Production inference platform for open-weights LLMs
OSChina's serverless hub for open model inference, tuning and apps
GPU marketplace and pay-per-token inference for 25+ open AI models
Pay-per-token inference API that Lambda itself is winding down
Acquired by NVIDIA in 2025, now DGX Cloud Lepton.
200+ models and per-second GPU cloud, priced low.
One API key, one bill, 300-plus LLMs
Brokered GPU inference: pay per token, skip the long-term contract
China's largest independent AI cloud, now filing for a Hong Kong IPO
Run any open AI model via API — no infrastructure
Open models at record tokens-per-second on RDU silicon
OpenAI-compatible API access to 200+ open models, billed per token
High-performance inference for open-weights LLMs
ByteDance's MaaS platform for Doubao, DeepSeek, GLM and Kimi
Baseten is in talks to raise US$1 billion at an US$11 billion valuation as inference money keeps flowing
Baseten, which rents Nvidia servers to companies running AI models, is in talks to raise US$1 billion at an US...
AI
1 Jun
Lablup open-sources MLXcel, an Apple-Silicon inference engine, under Apache 2.0
Lablup has released MLXcel, an open-source engine for running AI models on Apple Silicon, under the permissive...
KE
1 Jun