Overview
DeepMind's video-to-audio (V2A) technology generates rich soundscapes for videos using video pixels and natural language text prompts, opening up new creative opportunities for filmmakers and content creators.
Synchronized Audio Generation
V2A generates audio that is synchronized with the video input, creating a more immersive experience.
Unlimited Soundtrack Generation
V2A can generate multiple soundtracks for a video, allowing users to experiment with different audio outputs and choose the best match.
Guided Audio Generation
V2A allows users to define positive or negative prompts to guide the generated audio output, giving them more control over the final result.
Improved Lip Synchronization
V2A attempts to generate speech from input transcripts and synchronize it with characters' lip movements, creating a more realistic experience.
Get started
- Open the official website and confirm the service is available in your region.
- Check the current plan, usage limits and terms for your intended use.
- Try a small task with sample data before committing to a paid plan.
Editorial note
Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.
Record updated: Sep 20, 2026
Suggest a correction