Back to tools

SoundStorm

SoundStorm is a model for efficient, non-autoregressive audio generation. It receives as input the semantic tokens of AudioLM and relies on bidirectional attention and confidence-based parallel decoding to generate the tokens of a neural audio codec.

ActiveFirst listed: Oct 4, 2024
Visit website

Best for

Users seeking ai music generator solutions for their creative or business workflows.

Things to know

Usage limits, quotas, and premium capabilities are determined by the provider.

About this tool

Overview

SoundStorm is a model for efficient, non-autoregressive audio generation that produces high-quality audio two orders of magnitude faster than traditional autoregressive generation approaches.

Efficient Audio Generation

SoundStorm generates high-quality audio two orders of magnitude faster than traditional autoregressive generation approaches.

Dialogue Synthesis

SoundStorm can be used for dialogue synthesis by coupling it with the text-to-semantic modeling stage of SPEAR-TTS, allowing for the synthesis of high-quality, natural dialogues.

Detectable by Classifiers

SoundStorm-generated audio remains detectable by a dedicated classifier, with a detection rate of 98.5% using the same classifier as Borsos et al. (2022).

Non-Autoregressive Generation

SoundStorm uses bidirectional attention and confidence-based parallel decoding to generate the tokens of a neural audio codec.

Get started

  1. Open the official website and confirm the service is available in your region.
  2. Check the current plan, usage limits and terms for your intended use.
  3. Try a small task with sample data before committing to a paid plan.

Editorial note

Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.

Record updated: Sep 20, 2026

Suggest a correction