Wan 2.7 Image
One model that generates and edits — and actually gets the text right.
Wan 2.7 Image folds text-to-image and instruction editing into a single model, takes up to 9 reference images at once, and can run a reasoning pass before it draws. Below: what each of those actually does, where the resolution ceilings really sit, and what its text rendering did when we tested it.
This tool runs the standard Wan 2.7 Image tier at up to 2K, with Thinking Mode off for faster turnarounds. 4K output is exclusive to the separate Pro tier and to text-to-image. Every example image on this page was generated with this tool — they are real renders, not official marketing assets.
What Wan 2.7 Image actually changes
Three things separate Wan 2.7 Image from the Wan models before it. Each one comes with a limit worth knowing before you plan work around it.
Generate and edit without switching models
Earlier Wan releases split generation and editing across separate endpoints. Wan 2.7 Image handles both. Send a prompt alone and it generates; send the same prompt with up to 9 reference images and it edits, restyles, or fuses them. That matters for consistency work — keeping one character across a set, or holding a product steady while the scene around it changes.
The limit: the moment you attach a reference image, output caps at 2K on both tiers. 4K is text-to-image only. If you need a 4K composite, generate at 4K first and edit second — not the other way around.
A reasoning pass before the first pixel
Thinking Mode gives the model a planning step before it renders. It earns its keep when a prompt carries several constraints at once — three people each doing something specific, a sign whose wording has to survive, a layout described in sentences rather than shown. On a simple portrait or a loose landscape it changes little and costs you time.
The limit: it applies to text-to-image only — reference-guided edits ignore it — and every generation gets slower. The generator above keeps it off by default so results come back quickly.
Where the text holds up — and where it breaks
Alibaba lists 12-language text rendering. Rather than repeat that number, we ran it. Both images below came out of the generator on this page, first attempt, no retries and no cherry-picking.

Latin script: clean
Both the painted fascia and the hand-lettered chalkboard came back exactly as prompted, down to the em dash. Serif letterforms stayed consistent across the whole sign — no melted glyphs, no invented characters.

CJK: two out of three
Japanese (光と物質) and Simplified Chinese (光与物质) rendered correctly. Korean did not — the model returned 뭊과 물실 where 빛과 물질 was asked for, two wrong syllable blocks out of four.
The practical read: trust Wan 2.7 Image with Latin signage, Japanese, and Simplified Chinese. For Korean and other scripts, generate and then check the characters yourself — the failure is quiet, and a wrong syllable block looks perfectly confident inside an otherwise clean poster.
Several subjects, each doing its own thing
The other place Wan 2.7 Image pulls ahead is scenes where each subject has its own instruction. We asked for three cooks: one plating fish with tweezers, one wiping a rim, one calling out over a ticket rail. All three actions landed, faces stayed distinct, and the hands held together — historically the first thing to fall apart in a busy frame.

Wan 2.7 Image vs Wan 2.5 vs Wan 2.2
The three live Wan models solve different problems. Picking the wrong one wastes more time than picking the wrong prompt.
| Wan 2.7 Image | Wan 2.5 Image | Wan 2.2 | |
|---|---|---|---|
| Primary design | Unified generate + edit | Multimodal image generation | Video-first |
| Max resolution | 2K standard · 4K on Pro (text-to-image only) | Up to 2K | Video frames |
| Reference images | Up to 9 | Single reference | Single image for image-to-video |
| Reasoning pass | Thinking Mode (text-to-image) | Not available | Not available |
| Best used for | Text-in-image, multi-subject scenes, consistent sets | General stills, portraits | Short clips from text or a still |
Wan 2.7 Image parameters
The inputs that shape a Wan 2.7 Image request, and what each one costs you elsewhere.
promptstringThe description. Wan 2.7 rewards long, specific prompts — up to 5,000 characters.
aspect_ratioenum1:1, 16:9, 9:16, 4:3, 3:4, or 21:9. Omit it to let the model choose.
resolutionenum1K or 2K on the standard tier. 4K requires the Pro tier and text-to-image.
input_urlsstring[]Up to 9 reference images. Supplying any of these caps output at 2K.
thinking_modebooleanRuns the reasoning pass before generating. Text-to-image only; slower per image.
nintegerHow many images to return from one request.
How to use Wan 2.7 Image
Write it out in full
Name the subject, the light, the camera. If text should appear in the image, quote it exactly — the model renders what you put in quotes.
Pick a ratio, generate
Six ratios are available, from 1:1 to 21:9. The standard tier renders at up to 2K, which is enough for most web and print-preview work.
Refine or bring references
Change one detail per attempt so you can tell what moved. For consistency across images, switch to editing and upload up to 9 references.
Other models on this site
Wan 2.7 Image FAQ
What is Wan 2.7 Image?
Wan 2.7 Image is Alibaba’s unified image model, released April 1, 2026. One model handles both text-to-image generation and instruction-based editing, and it accepts up to 9 reference images in a single request. The standard tier renders up to 2K (2048×2048); the separate Pro tier reaches 4K on text-to-image only.
Is Wan 2.7 Image free to use here?
Yes — you can generate with Wan 2.7 Image on this page without paying upfront. New accounts get free credits, and each generation spends from that balance. No card is required to try it.
What is Thinking Mode in Wan 2.7 Image?
Thinking Mode is an extra reasoning pass the model runs before it starts drawing. It mainly helps on prompts with many constraints at once — several subjects doing specific things, embedded text that has to read correctly, or a layout described in words. It applies to text-to-image only, not to reference-image editing, and it makes each generation slower.
Can Wan 2.7 Image render text inside the picture?
Yes, and it is one of the model’s stronger areas. In our own testing, Latin-script signage came out clean and legible, including hand-lettered chalkboard text. Japanese and Simplified Chinese rendered correctly too. Korean was not reliable — a Hangul line came back with two wrong syllable blocks. Treat non-Latin text as something to verify rather than assume.
What is the difference between Wan 2.7 Image and Wan 2.7 Image Pro?
They are the same model family at different output ceilings. Standard tops out at 2K (2048×2048). Pro reaches 4K (4096×4096), but only for text-to-image — the moment you supply a reference image, both tiers cap at 2K. The generator on this page runs the standard tier.
How many reference images can Wan 2.7 Image take?
Up to 9 in one request. That is what makes it useful for keeping a character or product consistent across a set, fusing elements from several sources, or transferring a style while preserving a subject.
How does Wan 2.7 Image compare to Wan 2.5 and Wan 2.2?
Wan 2.2 is a video-first model. Wan 2.5 brought a stronger multimodal image path. Wan 2.7 is the first in the family built around a unified generate-and-edit design, with the reference-image ceiling raised to 9, a reasoning pass available on text-to-image, and a Pro tier that reaches 4K.
Which aspect ratios does the generator on this page support?
Six: 1:1, 16:9, 9:16, 4:3, 3:4, and 21:9. Wan 2.7 works from a ratio plus a resolution tier rather than arbitrary pixel dimensions, so pick the ratio that matches your output and let the model fill it.
Try Wan 2.7 Image free
Generate online with no card and no install. Free credits on signup, and the same model handles editing when you need it.
Open the generator