Overview
VLOGGER is a method for text and audio-driven talking human video generation from a single input image of a person, building on the success of recent generative diffusion models.
Text and Audio-Driven Generation
VLOGGER generates talking human videos from text and audio inputs, allowing for control over the content and tone of the video.
Stochastic Human-to-3D-Motion Diffusion Model
VLOGGER uses a stochastic human-to-3D-motion diffusion model to generate intermediate body motion controls, responsible for gaze, facial expressions, and pose.
Temporal Image-to-Image Translation Model
VLOGGER uses a temporal image-to-image translation model to generate the corresponding frames, taking the predicted body controls and a reference image of a person.
Diverse Video Generation
VLOGGER generates a diverse distribution of videos of the original subject, with a significant amount of motion and realism.
Get started
- Open the official website and confirm the service is available in your region.
- Check the current plan, usage limits and terms for your intended use.
- Try a small task with sample data before committing to a paid plan.
Editorial note
Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.
Record updated: Sep 20, 2026
Suggest a correction