Natively multimodal
One architecture for text, image, video, and audio — introduced with the 2.5 generation.
The generation that made Wan multimodal — and where it went next.
Wan 2.5 unified text, image, video, and audio in one architecture and set a new bar for in-image text rendering. This page covers what 2.5 introduced — and lets you generate with Wan 2.7, the current generation of the same image line, at native 2K.
The live tool runs Wan 2.7 Image — the current generation of the same Wanxiang image line that Wan 2.5 belongs to. Alibaba serves 2.5 only through its own cloud; 2.7 supersedes it here with native 2K and multi-reference editing. Every showcase image below is a real render from this tool.
4 models
T2V · I2V · T2I · editing
Multilingual
Legible in-image text
10s · 1080p
Video with synced audio
Native 2K
On Wan 2.7, served here
Announced at the 2025 Yunqi Conference, Wan 2.5-Preview moved the Wanxiang family onto a natively multimodal stack. For still images that meant something concrete: video-grade scene understanding applied to single frames — better depth, better lighting continuity, and text that actually reads.
Real renders · current Wan model
Every image below was generated with the tool at the top of this page — the Wan 2.7 successor to 2.5 — from a single prompt each. Cinematic light, frozen action, and one-pass typography are family traits that started with 2.5.

Dawn light, visible breath, drifting snow — the video stack's lighting sense applied to one still.

Headline, subtitle, and ornament border rendered in a single generation — the text strength Wan 2.5 was known for.

A leaping fox at 1/2000s with crisp fur against soft bokeh — photoreal detail without a real camera.
One architecture for text, image, video, and audio — introduced with the 2.5 generation.
Multilingual in-image typography for posters, packaging, and branded mockups.
Depth structure and lighting continuity inherited from the video stack.
Native 2K, finer detail, and up-to-9-reference editing in the generation served here.
Subject, lighting, mood, lens. Cinematic language — “dawn light, anamorphic, drifting snow” — plays to the family’s strengths.
For posters, put the headline in quotes and say where it goes. One clear instruction per text element keeps typography clean.
Render at native 2K, then adjust one variable at a time — or switch to image-to-image with up to 9 references.
Where the 2.5 generation sits between the open-weights 2.2 release and the current 2.7 image model.
| Generation | Wan 2.5 | Wan 2.7 (here) | Wan 2.2 |
|---|---|---|---|
| Released | Sep 2025 (preview) | 2026 · current | Jul 2025 |
| Image output | ~1080p class | Native 2K | — |
| Video output | 10s · 1080p · audio | — | 5s · 720p open model |
| Text rendering | Strong, multilingual | Strong, multilingual | Basic |
| Image editing | Single reference | Up to 9 references | — |
| Open weights | No (hosted) | No (hosted) | Yes |
| On this site | Superseded by 2.7 | 6 credits / image | Video generator |
Video · open model
The open-weights Wanxiang release — run text-to-video and image-to-video on the same family.
Open4 credits / image
Alibaba’s layout specialist: dense text, infographics, and multilingual typography.
Open2 credits / image
The fastest, cheapest drafting model here — 8-step generation, open source.
OpenWan 2.5 (Wan2.5-Preview) is the generation of Alibaba’s Tongyi Wanxiang model family announced at the Yunqi Conference on September 24, 2025. It shipped as four models — text-to-video, image-to-video, text-to-image, and image editing — on a natively multimodal architecture that handles text, images, video, and audio in one stack.
Wan 2.5’s still-image generator inherited the video stack’s scene understanding, which shows in its depth structure and prompt following. It became known for legible multilingual text rendering and “video-grade” lighting consistency — production-ready stills for product mockups and branded content.
The generator on this page runs Wan 2.7 Image — the current generation of the same Wanxiang image line, with native 2K output. Alibaba serves Wan 2.5 through its own cloud APIs; on this site you get the newer 2.7 model, which supersedes 2.5 on detail, resolution, and reference editing. The showcase images below are real Wan 2.7 renders from this tool.
Wan 2.7 pushes the image branch to native 2K output, sharper micro-detail, stronger prompt adherence, and multi-reference image-to-image editing (up to 9 input images). If you liked Wan 2.5’s text rendering and cinematic look, 2.7 keeps both and raises the resolution ceiling.
No. Unlike Wan 2.2 — which released open weights — Wan 2.5 launched as a hosted preview through Alibaba Cloud’s Bailian platform and the Tongyi Wanxiang site. That is why sites offering “Wan 2.5 online” actually proxy Alibaba’s API or, like this page, serve a newer generation of the same family.
Yes — multilingual text rendering was one of Wan 2.5’s headline strengths, and it carries into Wan 2.7. Posters with clean headlines, packaging mockups with readable labels, and bilingual layouts all work well; the botanical poster in the showcase was rendered in one pass, typography included.
Wan 2.7 Image costs 6 credits per image on Wan27Image. New Google sign-ups get free credits to try it, and the embedded generator above uses the same allowance as the full editor.
The Wan 2.5 release was video-first: its text-to-video and image-to-video models generate up to 10-second 1080p clips with synchronized audio. On this site, the video generator runs the related Wan 2.2 open model for video work — see the Wan 2.2 tool page for that.
Everything Wan 2.5 was praised for, at native 2K — free credits to start, 6 credits per image.
Looking for Wan video generation? Try the Wan 2.2 video tool