SiliconFlow 硅基流动 (SiliconCloud)
OpenAI-compatible API access to 200+ open models, billed per token
Overview
SiliconFlow runs SiliconCloud, a pay-as-you-go API serving 200+ open-weight models (DeepSeek, Qwen, GLM, Llama) plus image, video and audio generation. Founded in Beijing in 2023 by ex-Microsoft Research scientist Yuan Jinhui, it targets developers who'd rather not run their own GPUs.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
What you can produce with SiliconFlow 硅基流动 (SiliconCloud)
- OpenAI-compatible REST API across 200+ open-weight models
- Per-token pay-as-you-go billing, no subscription required
- Custom inference engine for accelerated low-latency serving
- Text, image, video and audio model endpoints
- Fine-tuning and custom/BYOC deployment for enterprise
- Free signup credit plus always-free small models (e.g. Qwen2.5-7B)
- Separate China (siliconflow.cn) and global (siliconflow.com) platforms
ASEAN Perspective
SiliconFlow 硅基流动 (SiliconCloud) in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
SiliconFlow's pitch is straightforward: one OpenAI-compatible endpoint, 200-plus open-weight models, priced by the token with no subscription lock-in. That combination made it a default pick for teams running DeepSeek, Qwen or GLM workloads without managing their own GPU fleet, and its June 2026 Series B (over 2 billion yuan, reportedly the largest MaaS round in China that year) suggests the growth is real. Free credits and always-free small models (Qwen2.5-7B) make it easy to trial.
The catch is verifiability: there's no G2, Capterra or Trustpilot presence to check claims against, and support is developer-forum-level rather than enterprise SLA. Rate limits scale with how much you've spent, which is fine for production but awkward for bursty free-tier testing. Solid default for cost-sensitive inference; less proven for teams needing contractual uptime guarantees or ASEAN data residency.
What people say
SiliconFlow doesn't show up on G2, Capterra, TrustRadius or Product Hunt — a Chinese-origin API provider serving a developer audience that mostly reviews products in Discord and GitHub issues rather than SaaS review sites. What there is points the same direction: developers cite the OpenAI-compatible endpoint as trivial to swap into existing tooling, and the pay-as-you-go model (no subscription, per-token billing across text, image, video and audio) as meaningfully cheaper than running comparable open models through providers like Together AI or Fireworks for Chinese-model-heavy workloads.
The company itself is young — founded in Beijing in August 2023 by Yuan Jinhui, a former Microsoft Research Asia principal researcher and founder of the OneFlow deep-learning framework, alongside co-founder Pan Yang. It raised an angel round that December and has since closed seven rounds; the June 2026 Series B, over 2 billion yuan, is described as the largest single raise in China's third-party Model-as-a-Service sector to date, putting the company's valuation near 7.74 billion yuan. That's a meaningful signal for a two-and-a-half-year-old infrastructure vendor.
Free tier is a genuine try-before-you-buy: new accounts get a starting credit, and some smaller models (Qwen2.5-7B among them) stay free indefinitely. Rate limits are tied to spend — accounts that have topped up at least a modest amount get materially higher request ceilings than unfunded free accounts, a reasonable anti-abuse mechanism, though it means testing at scale requires putting money down first.
What's missing from the public record is much independent complaint or praise beyond developer-forum chatter — no detailed outage post-mortems, no structured user reviews breaking down support quality. A GitHub issue on the Dify integration flags the need to self-limit embedding request rates to avoid bans, a minor but real operational wrinkle. For a platform this size serving production inference traffic, the lack of a visible support/SLA track record is the main open question for teams evaluating it beyond hobby or prototype use.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including SiliconFlow 硅基流动 (SiliconCloud)'s official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with SiliconFlow 硅基流动 (SiliconCloud) unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to SiliconFlow 硅基流动 (SiliconCloud) directly →
Spotted something out of date? Suggest an update →
Alternatives to SiliconFlow 硅基流动 (SiliconCloud)
More in LLMs & Chat