Tongyi Wanxiang 通义万相

Alibaba's open-weight image/video generator — now on Wan 2.7

Image Generation Freemium Has API
Researched · Published · Reviewed
RECATOOLS Score
6.8 / 10
Capability
7
Value for money
6
Ease of use
6
ASEAN readiness
6
API quality
7
Founded
HQ
Users
Hundreds of thousands of monthly downloads across HuggingFace/ModelScope variants
Launched
Originally China enterprise-only
Developer

Overview

Tongyi Wanxiang is Alibaba Cloud's text-to-image and video model family (Wan 2.1 through April 2026's Wan 2.7), with open Apache 2.0 weights plus a hosted DashScope API. It spans styles from watercolour to anime in Chinese and English, billed pay-as-you-go, not subscription.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

What you can produce with Tongyi Wanxiang 通义万相

  • A 4K print-ready image (up to 4096x4096) generated from a text prompt, with precise HEX brand colours locked and multilingual text rendered accurately in up to 12 languages (Wan 2.7 cloud/API only) — including Chinese, Arabic, and Korean
  • A 12-panel storyboard sequence with consistent character identity across all panels, produced from a set of reference images using Wan 2.7's multi-reference fusion and batch generation
  • A self-hosted, commercially deployable generation pipeline using the open Apache 2.0 weights (Wan 2.1/2.2 only — Wan 2.7 has no open weights); the 1.3B text-to-video model needs ~8.19 GB VRAM, so an 8 GB RTX 3070 is borderline and requires VRAM offloading/optimization — zero per-image API fees
  • A short narrative video clip (up to 15 seconds at 1080p) with synchronised audio, subject-voice cloning, and first/last frame control for cinematic scene transitions
  • Style-transferred marketing visuals in traditional Chinese ink-wash, retro, or 3D cartoon aesthetics generated from a product sketch or reference photo
  • An instruction-edited version of an existing video asset — changing background, lighting, or object appearance via text command without re-generating from scratch
  • Batch-generated consistent character assets (up to 12 images per run) for a game, animation, or brand campaign, maintaining skeletal pose and identity across varied angles and lighting
Advertisement

ASEAN Perspective

Tongyi Wanxiang 通义万相 in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Tongyi Wanxiang has grown well past a text-to-image tool: Wan 2.1 topped VBench, the open-weight Wan 2.1/2.2 releases (Apache 2.0) are genuinely popular in the open-source community, and April 2026's Wan 2.7 adds a "Thinking Mode" that plans a composition before rendering it. Backed by Alibaba Cloud's infrastructure and a documented DashScope API, it's a sensible pick for businesses already in that ecosystem, and the open weights give technical teams an option closed rivals don't.

The catch is access and polish: the consumer web app leans Mandarin-first, Wan 2.7's newest features are cloud/API-only with no open weights yet, and pay-as-you-go billing adds onboarding overhead compared to flat subscriptions. Clip length still caps around 15 seconds against competitors offering multi-minute output.

Independent AI-assisted assessment by RECATOOLS.

What people say

VBench doesn't lie about raw capability: Wan 2.1 topped the benchmark at 86.22%, and the open Apache 2.0 weights for Wan 2.1 and 2.2 have pulled hundreds of thousands of monthly downloads across HuggingFace and ModelScope — a rare case of a Chinese model the global open-source crowd actually adopted rather than just benchmarked. ChatForest rates it 4/5, and reviewers consistently praise the permissive licensing, non-expiring credits, and storyboard-level structural controls most closed competitors don't expose.

Wan 2.7, released in April 2026, is a different animal: cloud/API-only so far, priced around $0.10/second, and built around a "Thinking Mode" that plans composition before generating rather than reacting straight to the prompt. Its open weights haven't shipped yet, so anyone wanting the new 12-language text rendering or Thinking Mode locally is out of luck for now — that capability lives behind Alibaba Cloud's DashScope API only.

The consistent gripes: a steep learning curve on the structural controls, clips capped around 15 seconds against Kling's multi-minute output, physics simulation that still trails Runway Gen-4, and slow inference on consumer GPUs even with the open weights downloaded locally. Billing is pay-as-you-go rather than subscription — roughly ¥0.2 per image and ¥0.6-1 per second of video — which suits high-volume or API-first users better than someone paying per generation.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Tongyi Wanxiang 通义万相's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Tongyi Wanxiang 通义万相 unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Tongyi Wanxiang 通义万相 directly →

Spotted something out of date? Suggest an update →

Advertisement