AudioCraft
Transform text into stunning audio and soundscapes with ease.. [Free]
Last verified:
What is AudioCraft?
AudioCraft is Meta’s single-stop code base for generative audio, focused on music, sound effects, and audio compression from raw audio signals. It is positioned as AI research for audio and is built around controllable, high-quality text-to-audio generation. The website presents it as a unified framework for both creation and compression, rather than a narrow single-purpose generator.
The core system centers on three models: MusicGen, AudioGen, and EnCodec. MusicGen generates music from text and can also use melodic features for more control, while AudioGen generates environmental sounds from text prompts. EnCodec is the codec foundation that compresses audio into discrete tokens and reconstructs it with high fidelity, helping the generation pipeline work efficiently.
The site emphasizes that AudioCraft uses a simple autoregressive language-model approach over compressed audio tokens. It highlights long-term consistency, high-quality generation, and a streamlined model design compared with prior audio-generation work. The demos and model overview suggest it is aimed at researchers, developers, and creators who want to experiment with text-to-music or text-to-sound systems.
AudioCraft also presents itself as an open-source codebase for training and inference. That makes it especially relevant for people who want to study, extend, or build their own audio generative models rather than only use a closed online app. The website frames it as research infrastructure for interactive AI systems that help people co-create audio with models.
AudioCraft pricing
Pricing model: Freemium
The website does not list any paid plans or subscription tiers. It presents AudioCraft as open source and as a codebase for training and inference, so the practical cost is the user’s own infrastructure and implementation work rather than a posted product fee. No free tier, premium tier, or included plan limits are described on the site.
AudioCraft pros
- Single codebase for audio generation
- Covers music, sound effects, and compression
- Text-to-music generation via MusicGen
- Text-to-sound generation via AudioGen
- Uses EnCodec for high-fidelity compression
- Open-source training and inference code
- Supports controllable generation
- Supports text and melodic conditioning
- Designed for long-term audio consistency
- Built with a simple model architecture
- Uses autoregressive token-based generation
- Unified framework for research and experimentation
- Suitable for both developers and researchers
- Includes public demos and samples
- Supports raw-audio-based workflows
- Aims for high-quality output
- Can reconstruct audio from compressed tokens
AudioCraft cons
- Primarily a research codebase, not a consumer app
- No clear pricing page on the website
- Requires technical setup to use the code
- Not presented as a full editing suite
- Not focused on voice cloning
- Generation is limited to the provided model types
- Text prompts may need careful phrasing
- Best suited to audio research workflows
- Website gives limited product-style onboarding
- No explicit enterprise support details
- No subscription tiers listed
- Output control is narrower than a full DAW
- Sound generation is not framed as real-time performance
- Model quality depends on prompt and conditioning
- Licensing and deployment details are not fully spelled out
- Music and sound tools are separated into different models
- Not designed as a general media production platform
Frequently asked questions about AudioCraft
What is AudioCraft?
AudioCraft is Meta’s open-source codebase for generative audio research. It focuses on creating music, sound effects, and audio compression workflows from raw audio signals.
What models are included in AudioCraft?
AudioCraft consists of three main models: MusicGen, AudioGen, and EnCodec. MusicGen handles text-to-music generation, AudioGen handles text-to-sound generation, and EnCodec handles high-fidelity audio compression and reconstruction.
What does MusicGen do?
MusicGen generates music from text prompts and can also use melodic conditioning for more control. The website describes it as producing diverse and long music samples from user-provided text inputs.
What does AudioGen do?
AudioGen is designed for text-to-sound generation. It creates environmental and scene-based audio from a textual description, including realistic recording conditions and complex context.
What is EnCodec used for?
EnCodec is the audio codec foundation used in AudioCraft. It compresses audio into discrete tokens and reconstructs the waveform with high fidelity, supporting the generation pipeline for MusicGen and AudioGen.
Is AudioCraft open source?
Yes. The website describes AudioCraft as an open-source codebase for training and inference of generative audio models, intended to support research and innovation.
Who is AudioCraft for?
AudioCraft is aimed at researchers, developers, and technically oriented creators who want to study or build generative audio systems. It is especially relevant for people working on text-to-music, text-to-sound, or audio compression research.
Does AudioCraft support text prompts?
Yes. The site says AudioCraft supports music and audio generation from text inputs, with text-to-music through MusicGen and text-to-sound through AudioGen.
What kind of audio can it generate?
The website shows two main generation tasks: music generation and sound generation. It can create music samples and environmental sounds such as acoustic scenes, while highlighting high-quality and controllable outputs.
Does the website show pricing plans?
No. The website does not provide a pricing table, paid plans, or tiered subscriptions. It presents AudioCraft as a research codebase and open-source system rather than a commercial SaaS product.