Alibaba's unified image model · released April 1, 2026

Wan 2.7 Image

One model that generates and edits — and actually gets the text right.

Wan 2.7 Image folds text-to-image and instruction editing into a single model, takes up to 9 reference images at once, and can run a reasoning pass before it draws. Below: what each of those actually does, where the resolution ceilings really sit, and what its text rendering did when we tested it.

This tool runs the standard Wan 2.7 Image tier at up to 2K, with Thinking Mode off for faster turnarounds. 4K output is exclusive to the separate Pro tier and to text-to-image. Every example image on this page was generated with this tool — they are real renders, not official marketing assets.

2K → 4K
Standard, then Pro
9 images
References per request
5,000 chars
Prompt length accepted
Thinking Mode
Reasoning before drawing

What Wan 2.7 Image actually changes

Three things separate Wan 2.7 Image from the Wan models before it. Each one comes with a limit worth knowing before you plan work around it.

One model, two jobs

Generate and edit without switching models

Earlier Wan releases split generation and editing across separate endpoints. Wan 2.7 Image handles both. Send a prompt alone and it generates; send the same prompt with up to 9 reference images and it edits, restyles, or fuses them. That matters for consistency work — keeping one character across a set, or holding a product steady while the scene around it changes.

The limit: the moment you attach a reference image, output caps at 2K on both tiers. 4K is text-to-image only. If you need a 4K composite, generate at 4K first and edit second — not the other way around.

Thinking Mode

A reasoning pass before the first pixel

Thinking Mode gives the model a planning step before it renders. It earns its keep when a prompt carries several constraints at once — three people each doing something specific, a sign whose wording has to survive, a layout described in sentences rather than shown. On a simple portrait or a loose landscape it changes little and costs you time.

The limit: it applies to text-to-image only — reference-guided edits ignore it — and every generation gets slower. The generator above keeps it off by default so results come back quickly.

Text rendering, tested

Where the text holds up — and where it breaks

Alibaba lists 12-language text rendering. Rather than repeat that number, we ran it. Both images below came out of the generator on this page, first attempt, no retries and no cherry-picking.

Wan 2.7 Image render of a rain-slicked bookshop at blue hour, its painted sign reading THE MARGIN NOTE and a chalkboard reading Poetry 20% off, today only

Latin script: clean

Both the painted fascia and the hand-lettered chalkboard came back exactly as prompted, down to the em dash. Serif letterforms stayed consistent across the whole sign — no melted glyphs, no invented characters.

Wan 2.7 Image render of a gallery poster reading LIGHT & MATTER with Japanese and Simplified Chinese translations below it

CJK: two out of three

Japanese (光と物質) and Simplified Chinese (光与物质) rendered correctly. Korean did not — the model returned 뭊과 물실 where 빛과 물질 was asked for, two wrong syllable blocks out of four.

The practical read: trust Wan 2.7 Image with Latin signage, Japanese, and Simplified Chinese. For Korean and other scripts, generate and then check the characters yourself — the failure is quiet, and a wrong syllable block looks perfectly confident inside an otherwise clean poster.

Prompt adherence

Several subjects, each doing its own thing

The other place Wan 2.7 Image pulls ahead is scenes where each subject has its own instruction. We asked for three cooks: one plating fish with tweezers, one wiping a rim, one calling out over a ticket rail. All three actions landed, faces stayed distinct, and the hands held together — historically the first thing to fall apart in a busy frame.

Wan 2.7 Image render of three cooks working a restaurant pass under brass heat lamps, one plating a fish with tweezers, one wiping a plate rim, one calling out beside an orange ticket rail
Generated with the Wan 2.7 Image tool on this page, 16:9 at 2K, Thinking Mode off. Small background text on the tickets is not legible — the model reserves its text fidelity for type it was explicitly asked to render.

Wan 2.7 Image vs Wan 2.5 vs Wan 2.2

The three live Wan models solve different problems. Picking the wrong one wastes more time than picking the wrong prompt.

 Wan 2.7 ImageWan 2.5 ImageWan 2.2
Primary designUnified generate + editMultimodal image generationVideo-first
Max resolution2K standard · 4K on Pro (text-to-image only)Up to 2KVideo frames
Reference imagesUp to 9Single referenceSingle image for image-to-video
Reasoning passThinking Mode (text-to-image)Not availableNot available
Best used forText-in-image, multi-subject scenes, consistent setsGeneral stills, portraitsShort clips from text or a still

Wan 2.7 Image parameters

The inputs that shape a Wan 2.7 Image request, and what each one costs you elsewhere.

promptstring

The description. Wan 2.7 rewards long, specific prompts — up to 5,000 characters.

aspect_ratioenum

1:1, 16:9, 9:16, 4:3, 3:4, or 21:9. Omit it to let the model choose.

resolutionenum

1K or 2K on the standard tier. 4K requires the Pro tier and text-to-image.

input_urlsstring[]

Up to 9 reference images. Supplying any of these caps output at 2K.

thinking_modeboolean

Runs the reasoning pass before generating. Text-to-image only; slower per image.

ninteger

How many images to return from one request.

How to use Wan 2.7 Image

01

Write it out in full

Name the subject, the light, the camera. If text should appear in the image, quote it exactly — the model renders what you put in quotes.

02

Pick a ratio, generate

Six ratios are available, from 1:1 to 21:9. The standard tier renders at up to 2K, which is enough for most web and print-preview work.

03

Refine or bring references

Change one detail per attempt so you can tell what moved. For consistency across images, switch to editing and upload up to 9 references.

Other models on this site

Wan 2.7 Image FAQ

What is Wan 2.7 Image?

Wan 2.7 Image is Alibaba’s unified image model, released April 1, 2026. One model handles both text-to-image generation and instruction-based editing, and it accepts up to 9 reference images in a single request. The standard tier renders up to 2K (2048×2048); the separate Pro tier reaches 4K on text-to-image only.

Is Wan 2.7 Image free to use here?

Yes — you can generate with Wan 2.7 Image on this page without paying upfront. New accounts get free credits, and each generation spends from that balance. No card is required to try it.

What is Thinking Mode in Wan 2.7 Image?

Thinking Mode is an extra reasoning pass the model runs before it starts drawing. It mainly helps on prompts with many constraints at once — several subjects doing specific things, embedded text that has to read correctly, or a layout described in words. It applies to text-to-image only, not to reference-image editing, and it makes each generation slower.

Can Wan 2.7 Image render text inside the picture?

Yes, and it is one of the model’s stronger areas. In our own testing, Latin-script signage came out clean and legible, including hand-lettered chalkboard text. Japanese and Simplified Chinese rendered correctly too. Korean was not reliable — a Hangul line came back with two wrong syllable blocks. Treat non-Latin text as something to verify rather than assume.

What is the difference between Wan 2.7 Image and Wan 2.7 Image Pro?

They are the same model family at different output ceilings. Standard tops out at 2K (2048×2048). Pro reaches 4K (4096×4096), but only for text-to-image — the moment you supply a reference image, both tiers cap at 2K. The generator on this page runs the standard tier.

How many reference images can Wan 2.7 Image take?

Up to 9 in one request. That is what makes it useful for keeping a character or product consistent across a set, fusing elements from several sources, or transferring a style while preserving a subject.

How does Wan 2.7 Image compare to Wan 2.5 and Wan 2.2?

Wan 2.2 is a video-first model. Wan 2.5 brought a stronger multimodal image path. Wan 2.7 is the first in the family built around a unified generate-and-edit design, with the reference-image ceiling raised to 9, a reasoning pass available on text-to-image, and a Pro tier that reaches 4K.

Which aspect ratios does the generator on this page support?

Six: 1:1, 16:9, 9:16, 4:3, 3:4, and 21:9. Wan 2.7 works from a ratio plus a resolution tier rather than arbitrary pixel dimensions, so pick the ratio that matches your output and let the model fill it.

Try Wan 2.7 Image free

Generate online with no card and no install. Free credits on signup, and the same model handles editing when you need it.

Open the generator