The model that brought native sound to Alibaba's Wan family. Wan 2.5 generates dialogue, effects, and music in the same pass as the picture — turn a prompt or a single image into a 1080p clip up to 10 seconds long, with synchronized audio and lip-sync.
Talking-head spots, sound-designed product clips, and multilingual social ads made with Wan 2.5.
Turn a prompt or an image into a video with sound using Alibaba Wan 2.5 in three steps.
To get started, select the Wan 2.5 AI video model as your starting point.
Write a prompt describing the scene, action, and any dialogue you want spoken. Add a reference image to animate if you're starting from a still, and let Wan 2.5 generate the matching sound alongside the picture.
Watch it back with sound, then tune the prompt, the dialogue, or the framing and run the next variant.
Wan 2.5 on HeyOz generates video and sound together in a single pass — dialogue, effects, and music synced to the picture, from a prompt or a single image, at 1080p and in the language you need.
Wan 2.5 was the first model in Alibaba's Wan family to generate synchronized audio natively — human voice, sound effects, and background music produced in the same pass as the video, rather than dubbed on afterwards. Write the line you want spoken into the prompt and it comes back delivered on screen, with the sound already in the cut.
Describe the shot in text, or upload one photo — your product on a counter, a location, a character — and Wan 2.5 puts it in motion. The same native-multimodal model handles text-to-video and image-to-video, so you can begin from whichever asset you already have.
Generate at up to 1080p at a cinematic 24fps, in the aspect ratio your channel needs — 16:9 for YouTube and display, 9:16 for Reels and TikTok, 1:1 for feed. Draft on a lighter tier while you test angles, then finish the winner at full resolution.
Wan 2.5 generates spoken delivery with lip-sync and supports multiple languages, so you can produce the same spot for different markets with the voice built into the video. Localize a script and get lip-synced delivery back, rather than subtitles pasted over the original cut.
Built around one idea: video and sound in a single generation, from the assets you already have.
Wan 2.5 is Alibaba's video model from the Tongyi Wanxiang family, unveiled on September 24, 2025 as a preview release. It uses a native multimodal architecture across text, image, video, and audio, and it was the first model in the Wan series to generate synchronized human voice, sound effects, and visuals together in a single pass. On HeyOz it's wired straight into your brand kit and ad formats.
Pick Wan 2.5 as your model, write a prompt describing the scene and any dialogue you want spoken, and add a reference image if you're starting from a still. Generate, watch it back with sound, adjust, and run the next variant.
That's what HeyOz is built for. Point it at your product photo or write the scene, and the spot comes back on-brand, with sound, and sized for Meta, TikTok, and more.
You can start creating on HeyOz for free. See our pricing page for what's included on each plan.
Wan 2.5 generates clips of roughly 5 to 10 seconds, with the synchronized audio carried through the whole clip.
A text prompt, and optionally a reference image to animate for image-to-video. Because the model handles audio natively, you can write the dialogue into the prompt and have it spoken on screen — you don't need to supply a separate voice track.
Up to 1080p at 24fps, with lighter tiers available for drafting. Generate at 1080p for the cut you intend to run, and use a lower tier while you're testing angles.
Wan 2.5 introduced native synchronized audio to the family and generates from text or a single image. Wan 2.6, released later in December 2025, adds reference-to-video — casting a real person from a short clip of them — and longer generations up to 15 seconds. Pick 2.5 when you want fast text or image-to-video with sound; pick 2.6 when you're casting a specific spokesperson. Both are on HeyOz.
Pick the right engine for the job — all in one place.
The right tool for every kind of ad.
Anyone can give you Wan 2.5. Only HeyOz turns it into your ad.
HeyOz reads your brand and keeps every asset on-color, on-voice, on-message.
Seedance, Veo, Kling, GPT Image, and more — no juggling subscriptions.
Sized and formatted for Meta and TikTok, straight out of the box.
Not just a model — templates, avatars, and an agent that does it all.
Start Now. No agency, no brief, no blank screen.