Wan 2.7
by Alibaba
Smooth motion and scenes that stay coherent, from a prompt alone.
What it is
Wan 2.7 is described as "the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence". Those three claims are about the same underlying problem: video models tend to fall apart over time, and this one is tuned against that.
It works from a prompt alone, which is worth saying because a good number of video models require a starting image. If your idea does not begin with a picture you already have, this is one of the places to start.
It runs at 720p and 1080p, at five seconds, and supports adding your own audio track rather than generating one.
What it is good at
Scene fidelity and visual coherence are its stated focus. The background does not quietly become a different room halfway through.
Text to video from nothing. Useful when the idea exists as a sentence rather than as a picture.
Enhanced smoothness is the headline claim, and it is the thing you notice first when it is missing.
Where it struggles
- Alibaba publishes no list of limitations for this model, so none is invented here.
- Audio is something you supply, not something it generates. For sound made with the picture, Veo, Kling or LTX are the ones to use.
- Five seconds per generation. For longer, build a script so the shots belong together.
How to prompt it
Describe one continuous action and let the scene stay still around it. Because coherence is what this model is tuned for, giving it one thing to track works better than a prompt with three simultaneous events.
An example that shows what that means
A single red kite climbing steadily against a pale sky, the camera tilting up to follow it, thin cloud moving slowly behind, the line taut and visible against the light, nothing else in frame.
What we offer here
Straight from the catalogue, so this list never goes stale. The price is the provider’s published rate with our markup on top, and it is on the button before you run anything.
Questions
Does it need a starting image?
No , it works from a prompt alone, which is the main reason to pick it when your idea does not begin with a picture.
Does it generate sound?
No. You can supply your own audio track. For audio generated with the picture, use Veo, Kling or LTX.
How long is a clip?
Five seconds. For anything longer, build a script so several shots share one style and read as one piece.