- AI 模型目录
- Veo 3
Veo 3
参数配置
生成结果
样例展示
A surfer riding a massive ocean wave at sunset, aerial drone perspective, slow motion spray of water catching golden light
规格参数
定价
credits per second of video
~20 credits for an 8-second 1080p clip
Veo 3 is Google DeepMind's third-generation text-to-video model and the first in the Veo series to reach general API availability. It represents a step-change in AI video quality, combining Google's deep investment in physical simulation, multimodal understanding, and generative audio into a single cohesive model. Veo 3 produces 1080p video at 24 fps with up to 8 seconds of temporally coherent footage — and crucially, it generates synchronised audio alongside the visuals without requiring any additional processing.
能力说明
The standout capability of Veo 3 is its native audio generation: ambient sound effects, speech, and music are synthesised in lockstep with the video frames, producing a fully immersive clip from a single prompt. Beyond audio, Veo 3 excels at physically plausible scenes involving water, fire, and particle effects — areas where competing models frequently produce unnatural artefacts. It handles a wide range of cinematic styles from documentary realism to stylised animation, and supports three aspect ratios natively including vertical 9:16 for mobile-first content.
最佳实践
Veo 3 responds exceptionally well to prompts that describe both visual and auditory elements. Try "a thunderstorm rolling over a mountain range, distant rumble of thunder, rain hitting pine needles" to leverage its audio synthesis capability. For physical scenes, specify material properties: "heavy glass sphere falling into a shallow pool of water, high-speed splash, slow motion". Keep prompts under 200 words — Veo 3 handles dense scene descriptions well, but extremely long prompts can dilute the signal for less salient details.
常见问题
What makes Veo 3 different from other text-to-video models?
Veo 3 is the first text-to-video model in the Veo family to be available via public API, and it introduces a landmark capability: native audio generation. Unlike competing models that produce silent video, Veo 3 can synthesise ambient sound, dialogue, and music that is tightly synchronised with the visual content — all from a single text prompt.
How accurate is Veo 3 at simulating real-world physics?
Veo 3 is built on Google DeepMind's physical world model research, giving it an unusually accurate understanding of fluid dynamics, rigid body collisions, and cloth simulation. Scenes involving water, falling objects, or fabric in motion look significantly more believable compared to models trained without physics priors.
Can I use Veo 3 for commercial projects?
Yes. Videos generated via the Genace API using Veo 3 may be used for commercial purposes subject to Genace terms of service. The model supports creative, advertising, and brand content workflows at 1080p resolution.
相关模型
Veo 4
Veo 4 by Google DeepMind is the next generation text-to-video model featuring unprecedented realism, up to 4K resolution, and 60-second video generation. Join the waitlist to be first to access it on Genace.
Luma Ray 3
Luma AI
Luma Ray 3 brings 3D scene understanding and photorealistic ray tracing to AI video, producing cinematic camera motion and natural light behaviour that no other text-to-video model can match. Join the Genace waitlist for early access.
Wan 2.5
Alibaba
Wan 2.5 by Alibaba is an open-weight text-to-video model that leads academic benchmarks and supports local deployment. Generate high-quality video via the Genace API or explore self-hosting — join the waitlist for API early access.