Overview
MAGNeT is a novel approach to masked audio generation, leveraging a single non-autoregressive transformer to generate high-quality audio. This method operates directly on multiple streams of audio tokens, predicting spans of masked tokens during training and gradually constructing the output sequence during inference.
Single Non-Autoregressive Transformer
MAGNeT uses a single non-autoregressive transformer to generate high-quality audio, operating directly on multiple streams of audio tokens.
Masked Generative Sequence Modeling
MAGNeT predicts spans of masked tokens during training and gradually constructs the output sequence during inference.
Novel Rescoring Method
MAGNeT uses a novel rescoring method to enhance generated audio quality, leveraging an external pre-trained model to rescore and rank predictions.
Hybrid Version
MAGNeT offers a hybrid version that combines autoregressive and non-autoregressive modeling, allowing for flexible and efficient audio generation.
Get started
- Open the official website and confirm the service is available in your region.
- Check the current plan, usage limits and terms for your intended use.
- Try a small task with sample data before committing to a paid plan.
Editorial note
Pricing checked: Sep 20, 2026. Official website is active. Please verify current pricing and terms directly on the official site.
Record updated: Sep 20, 2026
Suggest a correction