Alibaba's Wan team launched Wan2.5-Preview today, marking only the second AI model capable of generating synchronized audio alongside video output. This breakthrough positions the model as a direct competitor to Google's Veo 3 in the race to solve one of generative AI's most stubborn challenges: creating audio that naturally matches visual content.
Key advancement: The model's unified framework handles multimodal generation natively, eliminating the need to sync separate audio and video streams post-generation.
Preview access: Currently available through a paid API while the team collects feedback for the full release, which will include open training and inference code.
Extended output: Doubles video length from 5 to 10 seconds with 1080p HD quality.
Frame-Perfect Sync: The audio alignment breakthrough that matters
The standout feature isn't just another video model—it's Wan2.5-Preview's ability to generate synchronized audio directly within the video creation process. Previous approaches required separate audio generation followed by complex alignment workflows, often resulting in mismatched timing or unnatural combinations.



