LogoGenace
  • Features
  • Models
  • Tools
  • Learn
  • Pricing
  • Docs
  1. AI Model Directory
  2. /
  3. Wan 2.5

Wan 2.5

Alibaba

Text to VideoComing Soonvideotext-to-videoalibabaopen-sourcecoming-soon

The try panel will be available when this model launches.

Coming Soon

This model is not yet available. Subscribe to get notified when it launches.

Specifications

720p1080pmax 10s16 fps16:99:161:1

Pricing

TBD — pricing not yet announced

Coming soon

Wan 2.5 is Alibaba's Tongyi Wanxiang (通义万相) team's contribution to the open-source AI video ecosystem — a fully open-weight text-to-video model that rivals proprietary alternatives on standard benchmarks. Released on Hugging Face with permissive weights, it represents the most capable open-source video model in the Alibaba research portfolio and has been widely adopted by the academic community for video generation research, fine-tuning experiments, and production deployments that require on-premises data sovereignty.

Capabilities

Wan 2.5 produces 720p and 1080p video at 16 fps across three aspect ratios, with clip durations up to 10 seconds. Its technical strengths lie in prompt adherence and scene composition — the model reliably places described subjects in the correct spatial arrangement with appropriate lighting and perspective. Because the weights are publicly available, the community has produced a rich ecosystem of fine-tuned variants specialised for specific domains such as anime style, product visualisation, and architectural walkthroughs. This ecosystem effect means the effective capability of Wan 2.5 extends far beyond the base model.

Best Practices

Wan 2.5 responds well to detailed, descriptive prompts that specify scene elements clearly. Because it was trained on diverse multilingual data, both Chinese and English prompts perform well — an advantage for international teams creating localised content. For best results at 720p, keep the scene composition relatively simple: one or two primary subjects with a described background. For 1080p output, you can introduce more compositional complexity, but specify foreground, midground, and background elements explicitly to help the model allocate detail budget appropriately.

Frequently Asked Questions

Is Wan 2.5 truly open source and can I deploy it locally?▾

Yes. Wan 2.5 is released with open weights under a permissive research licence on Hugging Face. The model can be deployed locally on consumer-grade GPUs (with quantised versions available for 24 GB VRAM setups), making it one of the few state-of-the-art video models accessible to independent researchers and developers without relying on a proprietary API.

How does Wan 2.5 compare to proprietary models on benchmarks?▾

Wan 2.5 achieved top scores on the EvalCrafter and VBench benchmarks at the time of its release, outperforming several closed-source models on metrics including motion smoothness, subject consistency, and prompt adherence. Its strong benchmark performance from an open-weight model makes it a significant milestone for the open AI video research community.

Why is Wan 2.5 capped at 16 fps rather than 24 fps?▾

Wan 2.5 targets a 16 fps output rate as a deliberate optimisation to balance motion quality with inference efficiency on accessible hardware. At 16 fps, the temporal resolution is sufficient for smooth-looking web and social media video, and the reduced frame count per clip makes local deployment viable on mid-range GPU hardware. The Alibaba team is actively developing higher-fps variants for future releases.

Related Models

Hunyuan Video
Coming Soon
Text to Video

Hunyuan Video

Tencent

Hunyuan Video by Tencent is an open-source video model that generated massive community interest at launch for its strong motion coherence and Chinese-language scene understanding. Join the Genace waitlist to access it via API.

videotext-to-videotencent
Try it →
Veo 3
Text to Video

Veo 3

Google

Generate stunning 1080p videos with Veo 3 by Google DeepMind. The first commercially available Veo model features native audio generation, physics-accurate simulation, and up to 8 seconds of cinematic video from text prompts.

videotext-to-videogoogle
Try it →
Luma Ray 3
Coming Soon
Text to Video

Luma Ray 3

Luma AI

Luma Ray 3 brings 3D scene understanding and photorealistic ray tracing to AI video, producing cinematic camera motion and natural light behaviour that no other text-to-video model can match. Join the Genace waitlist for early access.

videotext-to-videoluma
Try it →

Compare with Other Models

  • Wan 2.5 vs Veo 4
  • Wan 2.5 vs Veo 3
  • Wan 2.5 vs Sora 2
  • Wan 2.5 vs Luma Ray 3
LogoGenace

Multimodal generation gateway for AI agents. Video, image, one endpoint.

Product
  • Features
  • Pricing
  • FAQ
Resources
  • Documentation
  • Learn
  • Tools
  • Benchmark
  • Error Codes
  • Changelog
  • Roadmap
Company
  • About
  • Contact
  • Waitlist
Legal
  • Cookie Policy
  • Privacy Policy
  • Refund Policy
  • Terms of Service
© 2026 Genace. All Rights Reserved.