MusicCaps

MusicCaps is a dataset of 5.5k high-quality music captions written by musicians, useful for text-to-music generation research and automatic music captioning systems.

Last verified:

Visit MusicCaps

What is MusicCaps?

MusicCaps is a dataset released by Google AI alongside the MusicLM model, containing 5,521 music examples each labeled with an English aspect list and a free text caption written by musicians. The dataset provides high-quality music-text pairs with rich descriptions from human experts, supporting research in text-to-music generation and automatic music captioning.

Key features include 10-second audio clips in WAV format paired with detailed captions describing instruments, genres, mood, tempo, and musical characteristics. Each example includes an aspect list with structured tags and a free-text caption. The dataset is licensed under CC BY-SA 4.0, allowing commercial use with attribution and share-alike requirements. Audio files are sourced from YouTube and provided via links.

MusicCaps is designed for AI researchers, music technology developers, and machine learning practitioners working on text-to-music generation models, automatic music captioning systems, audio classification, and multimodal music understanding. It serves as a training corpus for models like MusicLM that generate high-fidelity music from text descriptions.

MusicCaps pricing

Pricing model: Free

Free - The MusicCaps dataset is publicly available for download on Kaggle and Hugging Face under CC BY-SA 4.0 license. No paid tiers exist as this is an open research dataset. Users can freely download and use the dataset for research and commercial purposes with attribution required.

MusicCaps pros

  • 5,521 music-text pairs provide substantial training data
  • Captions written by musicians ensure high-quality descriptions
  • Rich aspect lists with structured tags for each example
  • CC BY-SA 4.0 license allows commercial use
  • 24 kHz audio quality for high-fidelity generation training
  • Supports hierarchical sequence-to-sequence modeling research
  • Free text captions describe instruments, mood, tempo, and style
  • Enables text-and-melody conditioning research
  • Validated by MusicLM outperforming previous systems
  • Supports long music generation up to several minutes
  • Diverse genres including rock, jazz, classical, folk, electronic
  • Includes both instrumental and vocal music examples
  • Publicly available for research on Kaggle and Hugging Face
  • Enables story mode generation with sequential prompts
  • Supports painting-to-audio conditioning research

MusicCaps cons

  • Only 5,521 examples is small compared to modern datasets
  • Audio not directly provided, only YouTube links
  • 10-second clips are short for some applications
  • Some recordings have low or poor audio quality
  • YouTube copyright restrictions may limit commercial use
  • CC BY-SA 4.0 requires sharing derivatives under same license
  • No official API or download tool from Google
  • Dataset may contain biases from YouTube source material

Frequently asked questions about MusicCaps

What is MusicCaps?

MusicCaps is a dataset composed of 5.5k music-text pairs released by Google Research alongside the MusicLM model. Each of the 5,521 music examples is labeled with an English aspect list and a free text caption written by musicians, providing rich descriptions for text-to-music generation research.

What license does MusicCaps use?

MusicCaps is licensed under CC BY-SA 4.0 (Creative Commons Attribution-ShareAlike 4.0 International). This allows commercial use, distribution, and creation of derivative works, but requires attribution to Google and mandates that adaptations be shared under the same license.

How many music examples are in MusicCaps?

The dataset contains exactly 5,521 music examples, each paired with an aspect list and free text caption. All examples are 10-second audio clips in WAV format at 24 kHz quality.

Who wrote the captions in MusicCaps?

The captions were written by musicians and human experts, ensuring high-quality, accurate descriptions of the music. These expert-written captions describe instruments, genres, mood, tempo, and other musical characteristics.

Where can I download MusicCaps?

MusicCaps is available on Kaggle at https://www.kaggle.com/datasets/googleai/musiccaps and on Hugging Face at https://huggingface.co/datasets/google/MusicCaps. Unofficial download repositories exist on GitHub for extracting audio from YouTube links.

What is the audio quality in MusicCaps?

The dataset contains 10-second audio clips at 24 kHz quality. However, some examples have low quality, noisy, or poor audio as noted in the aspect lists, since the audio is sourced from YouTube videos with varying recording quality.

Can I use MusicCaps for commercial projects?

Yes, the CC BY-SA 4.0 license allows commercial use. However, you must provide attribution to Google as the originator and any derivative works must be shared under the same CC BY-SA 4.0 license. Original audio remains hosted on YouTube with its copyright.

What is MusicLM?

MusicLM is Google's AI model for generating high-fidelity music from text descriptions like 'a calming violin melody backed by a distorted guitar riff'. It uses hierarchical sequence-to-sequence modeling, generates music at 24 kHz consistent over several minutes, and can be conditioned on both text and melody.

What formats are available in MusicCaps?

The dataset provides audio in WAV format (extracted from YouTube) and annotations in JSON format. The data includes fields like ytid (YouTube ID), start_s, end_s, audioset_positive_labels, aspect_list, caption, author_id, and subset flags.

What research does MusicCaps support?

MusicCaps supports research in text-to-music generation, automatic music captioning, audio classification, multimodal music understanding, melody conditioning, story mode generation with sequential prompts, and painting-to-audio conditioning. It trained MusicLM which outperforms previous systems in audio quality and text adherence.

Categories

Use cases

Browse all AI tools on NeedAnAI