Zhipu Qingying 智谱清影 (CogVideoX)

Zhipu's free video generator, built on the open-weight CogVideoX model

Video & Audio Freemium Has API Open Source
Researched · Published · Reviewed
RECATOOLS Score
7.3 / 10
Capability
7
Value for money
8
Ease of use
6
ASEAN readiness
6
API quality
7
Founded
HQ
Users
Launched
Developer

Overview

Qingying (清影) is Zhipu AI's free consumer video generator for text-to-video and image-to-video, built on the open-source CogVideoX model developed with Tsinghua University. Businesses can call the same capability via API on Zhipu's bigmodel.cn open platform.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

What you can produce with Zhipu Qingying 智谱清影 (CogVideoX)

  • Free text-to-video and image-to-video generation (consumer app)
  • 10-second 1080p output with CogSound synced audio
  • Multi-shot generation and shot control
  • Open-source CogVideoX weights on GitHub (2B under Apache 2.0)
  • Metered developer API via bigmodel.cn (OpenAI-style integration)
  • Self-hostable on a single ~16GB VRAM consumer GPU
Advertisement

ASEAN Perspective

Zhipu Qingying 智谱清影 (CogVideoX) in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Qingying's real significance isn't the consumer app — it's that Zhipu released the underlying CogVideoX weights on GitHub under a permissive license (Apache 2.0 for the 2B variant) rather than keeping them closed like Sora or Kling. That makes it one of the only credible open-weight video models developers can actually self-host and fine-tune, and CogVideoX is regularly cited as the strongest option specifically for following detailed, multi-clause prompts, even where rivals like HunyuanVideo or Wan edge it out on raw visual polish.

The hosted consumer product is free, generating up to 10-second 1080p clips with synced audio via CogSound, but it's a Mandarin-first interface tied to the Zhipu Qingyan app and chatglm.cn. Against 2026's frontier video models (Kling v3, Sora, Seedance) Qingying's output quality is competitive but not category-leading. Best fit: developers who want an open, fine-tunable video model rather than a locked-down API.

Independent AI-assisted assessment by RECATOOLS.

What people say

Qingying (清影, "clear shadow") is Zhipu AI's video generation product, first launched in mid-2024 and now on its second major version. It runs on CogVideoX, a text/image-to-video model Zhipu built with a Tsinghua University research group — the same lineage as Zhipu's ChatGLM line, now rebranded internationally as Z.ai. The consumer product sits inside the Zhipu Qingyan chat app and at chatglm.cn/video, offering free text-to-video and image-to-video generation with no waitlist, which Chinese tech press at launch described as a rare "immediately usable, actually free" release compared to the invite-gated rollout of OpenAI's Sora at the time.

What sets Qingying apart from most consumer video tools is that the underlying model isn't locked away. Zhipu published CogVideoX on GitHub (now under the zai-org organization) with the 2B-parameter version under a permissive Apache 2.0 license and the larger 5B version under a separate, more restrictive CogVideoX license distributed via Hugging Face. That's made CogVideoX-5B one of the standard reference points in open-source video-model comparisons — reviewers consistently rank it as the strongest open-weight model specifically for accurately following detailed, multi-clause text prompts, even when competitors like HunyuanVideo (visual quality) or Wan 2.2 (versatility) score higher on other axes. It needs roughly 16GB of VRAM to run locally, accessible on a single high-end consumer GPU rather than requiring a server-grade card.

The Qingying v2.0 update pushed output to 10-second, 1080p clips with tighter motion and shot control, multi-shot generation, and CogSound for synced audio generation alongside video. Enterprise and developer access runs through the bigmodel.cn open platform as a metered API, with new accounts credited roughly ¥18 to test it; Zhipu hadn't published a flat per-video CogVideoX API rate on its public pricing page as of this research.

Where it sits competitively: by mid-2026, text-to-video arena leaderboards put newer frontier models — Kling v3, Seedance 2.0 — ahead of CogVideoX on raw output quality, so Qingying isn't the tool to reach for if you need the single best-looking clip. It's the tool to reach for if open weights, self-hosting or fine-tuning a video model matters more than topping a leaderboard.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Zhipu Qingying 智谱清影 (CogVideoX)'s official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Zhipu Qingying 智谱清影 (CogVideoX) unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Zhipu Qingying 智谱清影 (CogVideoX) directly →

Spotted something out of date? Suggest an update →

Advertisement