StepFun 阶跃星辰

Chinese multimodal frontier-model lab founded by Microsoft Research veteran

LLMs & Chat Freemium Has API
Researched · Published
RECATOOLS Score
6.6 / 10
Capability
7.5
Value for money
7
Ease of use
5
ASEAN readiness
4
API quality
6
Founded
2023
HQ
Shanghai, China
Users
Launched
Developer

Overview

StepFun is a Shanghai-based frontier-model lab founded by Jiang Daxin (former Microsoft Research Asia chief scientist). Publishes the Step series of LLMs — Step-1 (chat), Step-2 (frontier), Step-Video (text-to-video), Step-Audio — competing across the full multimodal stack against Baidu, Tencent and ByteDance.

---

阶跃星辰由前微软亚洲研究院首席科学家姜大昕创立,总部位于上海。发布 Step 系列模型——Step-1(对话)、Step-2(前沿)、Step-Video(文生视频)、Step-Audio——在多模态全栈层面与百度、腾讯、字节展开正面竞争。

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 19 May 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

Use cases

Multimodal chat Video generation Audio AI Enterprise LLM

What you can produce with StepFun 阶跃星辰

  • Call Step 3.5 Flash through the StepFun open platform API to power an agent with tool use and a 256k context window.
  • Download the Apache 2.0-licensed Step 3.5 Flash weights and self-host the model on your own inference infrastructure.
  • Generate short videos from Chinese or English text prompts using the open-sourced Step-Video-T2V model.
  • Build a voice assistant with speech recognition, emotional speech synthesis and voice cloning using the Step-Audio model family.
  • Chat with the consumer 阶跃AI app for Chinese-language question answering, writing and image understanding.
  • Fine-tune the open Step-Audio 2 mini model on your own audio data for domain-specific speech understanding.
  • Benchmark your speech pipeline against StepEval-Audio-360, StepFun's published multi-turn Chinese audio evaluation set.
Advertisement

ASEAN Perspective

StepFun 阶跃星辰 in Southeast Asia

团队学术背景与多模态全栈布局是其差异化。在中国市场被视为继 MiniMax、Moonshot、Zhipu 之后的下一档头部实验室。

RECATOOLS Verdict

StepFun (阶跃星辰) is a well-funded Chinese AI lab building capable multimodal Step models, with an OpenAI-compatible API, realtime voice tooling and recent open-source releases such as Step 3.7 Flash aimed at agentic use. For developers wanting access to a strong China-origin multimodal stack, the models are genuinely competitive.

The big caveat for international and ASEAN users: StepFun's English-facing web app (stepfun.ai) was discontinued in May 2026, leaving primarily the China-hosted platform (api.stepfun.com), which raises practical hurdles around documentation, billing, latency and data residency for non-China teams. Open-weight releases are the most accessible path outside China.

Independent AI-assisted assessment by RECATOOLS.

What people say

StepFun has grown from a well-funded 2023 startup into one of the more credible names among China's frontier-model labs, and as of 2026 it is very much alive: the Shanghai company closed a record domestic funding round in January 2026 with Megvii founder Yin Qi joining as chairman, has been reported to be raising a roughly $2.5 billion pre-IPO round, and is preparing a Hong Kong listing. Reported revenue of around 500 million yuan in 2025 came largely from smartphone partnerships — StepFun models ship on-device with OPPO, Honor and ZTE handsets, covering a meaningful share of China's leading phone makers.

Among developers, StepFun's reputation rests mostly on its open-weight releases rather than its consumer chat app. Step-Video-T2V and the 130B-parameter Step-Audio drew genuine attention on Hacker News and GitHub when they were open-sourced in early 2025, and the February 2026 release of Step 3.5 Flash — a 196-billion-parameter MoE model with 11 billion active parameters, a 256k context window and an Apache 2.0 licence — positioned it as a serious open alternative for agent workloads. Developers generally praise the permissive licensing and the audio stack in particular, which handles Chinese speech, voice cloning and emotion control better than most Western open models.

The caveats are the usual ones for Chinese labs: documentation and community support skew heavily toward Chinese, the consumer-facing 阶跃AI app is domestically oriented, and English-language benchmarks tend to trail the very top US frontier models. Content served through the official platform is also subject to Chinese regulatory filtering.

StepFun fits developers who want capable open-weight multimodal models — especially speech and video — under Apache licensing, and teams building for Chinese-language users. It is a weaker pick for English-first products or anyone needing mature Western-style enterprise support.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on

This entry was compiled from publicly available data including StepFun 阶跃星辰's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with StepFun 阶跃星辰 unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to StepFun 阶跃星辰 directly →

Spotted something out of date? Suggest an update →

Advertisement