Back to tools

Voicebox

Voicebox is a cutting-edge speech generative model built upon Meta’s non-autoregressive flow matching model. It outperforms single-purpose AI models across speech tasks through in-context learning, synthesizing speech across six languages, removing transient noise, editing content, transferring audio style within and across languages, and generating diverse speech samples.

ActiveFirst listed: Oct 4, 2024
Visit website

Best for

Users seeking ai speech synthesis tools for their daily workflows.

Things to know

Usage limits, quotas, and premium capabilities are determined by the provider.

About this tool

Overview

Introducing Voicebox, a revolutionary speech generative model that synthesizes speech across six languages, removes transient noise, edits content, transfers audio style, and generates diverse speech samples.

Multilingual Support

Synthesizes speech across six languages: English, French, German, Spanish, Polish, and Portuguese.

Transient Noise Removal

Removes transient noise by re-generating noise-corrupted speech.

Content Editing

Corrects misspoken words without having the speaker re-record the audio.

Audio Style Transfer

Transfers audio style within and across languages.

Get started

  1. Open the official website and confirm the service is available in your region.
  2. Check the current plan, usage limits and terms for your intended use.
  3. Try a small task with sample data before committing to a paid plan.

Editorial note

Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.

Record updated: Sep 20, 2026

Suggest a correction