Overview
SoundStorm is a model for efficient, non-autoregressive audio generation that produces high-quality audio two orders of magnitude faster than traditional autoregressive generation approaches.
Efficient Audio Generation
SoundStorm generates high-quality audio two orders of magnitude faster than traditional autoregressive generation approaches.
Dialogue Synthesis
SoundStorm can be used for dialogue synthesis by coupling it with the text-to-semantic modeling stage of SPEAR-TTS, allowing for the synthesis of high-quality, natural dialogues.
Detectable by Classifiers
SoundStorm-generated audio remains detectable by a dedicated classifier, with a detection rate of 98.5% using the same classifier as Borsos et al. (2022).
Non-Autoregressive Generation
SoundStorm uses bidirectional attention and confidence-based parallel decoding to generate the tokens of a neural audio codec.
Get started
- Open the official website and confirm the service is available in your region.
- Check the current plan, usage limits and terms for your intended use.
- Try a small task with sample data before committing to a paid plan.
Editorial note
Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.
Record updated: Sep 20, 2026
Suggest a correction