Text or image in. A complete audiovisual scene out.
Wan 2.5 AI Video Generator
Add Image
JPG, PNG, WebP
Max 10MB
The output video aspect ratio will match your uploaded image
Ready to Create
Configure your settings and click generate to start creating amazing videos
Wan 2.5 Video Examples with Native Audio
See how Wan 2.5 transforms text and images into complete audio-visual experiences
Image to Video with Audio
Transform static images into dynamic videos with synchronized soundtracks, speech, and environmental audio
Input

Text to Video with Native Audio
Create complete videos with visuals, speech, and music from text descriptions alone
Input
“A dimly lit jazz bar at night, wooden tables glowing under warm pendant lights. Patrons sip drinks and chat quietly while a three-piece band performs on stage. The saxophone player stands under a spotlight, gleaming instrument reflecting the light. No dialogue. Ambient audio: smooth live jazz music with saxophone and piano, clinking glasses, low murmur of audience conversations, occasional burst of laughter from a nearby table. Camera: slow pan across the crowd, then gentle zoom toward the saxophone player’s solo, focusing on expressive hand movements.”
Create complete scenes with Wan 2.5
Wan 2.5 combines visual generation and audio direction in one focused workflow for social clips, concepts, ads, and story moments.
Coordinated video and audio
Describe dialogue, ambience, music, and sound effects alongside the action. Wan 2.5 uses those cues to shape an audiovisual result without requiring a separate soundtrack workflow.
Text-to-video and image-to-video
Start from a prompt when you want a new composition, or upload an image when you need the generated motion to follow an existing subject, palette, or layout.
Flexible output controls
Choose a 5- or 10-second duration, 720p or 1080p resolution, and a landscape, portrait, or square aspect ratio where the selected generation mode supports it.
Prompt controls for more intent
Use a negative prompt to reduce unwanted elements and an optional seed when you want a repeatable starting point for experimentation.
How to generate a Wan 2.5 video
Move from an idea to a downloadable video in three straightforward steps.
1. Choose your starting point
Select text-to-video and describe the scene, or choose image-to-video and upload a clear reference image. Include subject motion, camera direction, setting, and audio cues in the prompt.
2. Set the output
Pick the duration, resolution, and aspect ratio that fit the destination. Add a negative prompt or seed only when you need more control.
3. Generate and review
Submit the task, preview the finished result, and download the video. If you iterate, change one prompt detail at a time so you can see what improves the scene.
Wan 2.5 questions
Practical answers about Wan 2.5 inputs, output settings, audio direction, prompting, and usage.
What is Wan 2.5?
Wan 2.5 is an AI video model for generating short videos from text prompts or reference images. This workspace supports 5- and 10-second output in 720p or 1080p, with audio cues included in the prompt.
Can Wan 2.5 generate audio with video?
Yes. You can describe dialogue, music, ambience, or sound effects in the same prompt as the visual scene. Clear, concrete audio instructions usually give the model a better target.
Which Wan 2.5 input modes are available?
Use text-to-video to create a scene from a description, or image-to-video to animate a reference image. Image-to-video is useful when composition or subject appearance matters.
What duration and resolution can I choose?
The workspace offers 5- or 10-second generation at 720p or 1080p. Higher resolution and longer duration use more credits and can take longer to process.
How should I write a Wan 2.5 prompt?
Name the subject, action, setting, camera movement, visual style, lighting, and desired sound. Keep the direction internally consistent and reserve the negative prompt for specific elements you want to avoid.
Can I use a generated video commercially?
Commercial use depends on your plan terms, the provider terms, and the rights attached to your prompts, uploads, music directions, brands, and depicted people. Review those requirements before publishing commercial work.
Why can two generations look different?
Generative video is probabilistic, so the same prompt can produce variations. Use a seed when available, keep the reference image unchanged, and adjust one instruction at a time for more controlled iteration.
Explore more AI Video
Switch tools without breaking your creative flow.
Turn your next idea into motion and sound
Choose text or an image, describe the scene clearly, and generate a short video ready to review and refine.
