Mixture of experts
Wan 2.2 increases model capacity by assigning high-noise and low-noise denoising stages to specialized experts.
Create short cinematic clips with Wan 2.2 online. Start from a written scene or animate a still image, then choose 480p or 720p output without downloading model weights or building a local ComfyUI workflow.
T2V + I2V
Two live workflows
480p / 720p
Hosted output
~5 seconds
Current clip length
A14B MoE
Fast hosted variants
LIVE TEXT-TO-VIDEO RESULT
Wan 2.2 works best when the prompt describes what changes over time. Name the subject, its action, the camera move, the lighting, and the details that should remain stable. The sample was generated during this page's backend verification, not borrowed from a model gallery.
PROMPT
A small red paper boat glides through a rain-slick city alley at blue hour, cinematic tracking shot, realistic water reflections, subtle wind moving paper edges.
WAN 2.2 IMAGE TO VIDEO
Upload an image and prompt the motion rather than restating every visual detail. Wan 2.2 follows the source aspect ratio and uses that single image as the visual starting point. The source below was created for this page, then animated by the live Wan 2.2 image-to-video endpoint.

MOTION PROMPT
The mechanical hummingbird unfolds its translucent wings, beats them rapidly, then lifts above the wet copper wire while the camera makes a gentle push-in. Preserve the rooftop garden, sunrise lighting, bird design, and realistic materials.
WHAT IS WAN 2.2?
Released on July 28, 2025, Wan 2.2 expanded the Wan video family with dedicated A14B text-to-video and image-to-video models plus a compact TI2V-5B hybrid. The official project describes a mixture-of-experts architecture that assigns different denoising stages to specialized experts.
Read the official Wan 2.2 repositoryWan 2.2 increases model capacity by assigning high-noise and low-noise denoising stages to specialized experts.
Training labels cover lighting, composition, contrast, and color tone, making visual direction easier to express in a prompt.
The release expanded image and video training data over Wan 2.1 to improve motion, semantics, and aesthetic range.
The hosted A14B variants expose both 480p and 720p, while the official TI2V-5B release targets 720p at 24 fps.
Searchers looking for a Wan 2.2 prompt usually need control, not extra adjectives. Build the shot in four layers and remove instructions that do not change the visible video.
01
Name the main subject and the stable visual details that identify it.
02
Describe the action as a sequence: starts, changes, then settles.
03
Choose one clear move such as a push-in, orbit, pan, or locked shot.
04
State what must stay stable and exclude text, logos, or scene cuts when needed.
1
Use Text to Video for a new scene. Use Image to Video when composition and subject identity already exist in a still frame.
2
Write the motion prompt, select 480p or 720p, and choose landscape or vertical orientation for text-to-video.
3
Review the five-second result, then change one variable at a time: motion strength, camera move, timing, or stability constraints.
“Wan 2.2” refers to a family, not one checkpoint. This page uses optimized A14B endpoints for convenient browser generation; the official releases remain useful when you need local control.
| Capability | Hosted Wan 2.2 | T2V-A14B | I2V-A14B | TI2V-5B |
|---|---|---|---|---|
| Text to video | Yes | Yes | No | Yes |
| Image to video | Yes | No | Yes | Yes |
| Resolution | 480p / 720p | 480p / 720p | 480p / 720p | 720p |
| Typical setup | Browser | Local GPU | Local GPU | Local GPU |
| Best fit | Fast online tests | Prompt-led scenes | Still-image motion | Lower-VRAM local use |
Wan 2.2 is an open video generation model family from the Wan team. It introduced a mixture-of-experts video diffusion architecture and includes separate text-to-video, image-to-video, and hybrid 5B releases.
Yes. The generator above connects to hosted Wan 2.2 text-to-video and image-to-video endpoints, so you do not need to download model weights or configure ComfyUI.
Yes. Upload a source image, describe the motion, and generate a short clip that follows the original composition and aspect ratio.
The hosted Wan 2.2 endpoints on this page support 480p and 720p. Text-to-video supports 16:9 and 9:16. Image-to-video follows the aspect ratio of the source image.
The current online workflow creates approximately five-second clips. This keeps generation practical for motion tests, social shots, concept frames, and short inserts.
No. The Wan 2.2 text-to-video and image-to-video endpoints used here generate silent video. Add music, dialogue, or sound effects in a separate audio or editing workflow.
The Wan team released model weights and inference code under the repository license. This page offers a hosted workflow for people who prefer not to run the models locally.
Describe the subject, the action over time, camera movement, environment, lighting, and constraints. For image-to-video, focus the prompt on motion and what must remain stable.
Test a prompt in the embedded tool, or open the full workspace when you need generation history, downloads, and a larger editing surface.
Working with still images instead? Try the GPT Image 2 generator