Harmonai

Harmonai.org is an open-source, generative audio tool developed by the Stability AI Lab. It aims to enhance music production accessibility ...

Last verified:

Visit Harmonai

What is Harmonai?

Harmonai is a Stability AI Lab that releases open-source generative audio tools to make music production more accessible and fun for everyone. Its flagship model, Dance Diffusion, uses machine learning diffusion models to generate unique audio waveforms from scratch, starting from random noise and refining it into music rather than relying on pre-existing samples or MIDI clips.

Key features include unconditional random audio sample generation, audio sample regeneration and style transfer using a single audio file, audio interpolation between two audio files, and the ability to create custom infinite sound libraries. The platform offers six pre-trained models trained on different datasets including piano recordings, glitch sounds, wildlife recordings, and daily song projects. Users can fine-tune models with custom datasets and adjust parameters like tempo, pitch, volume, and style.

Harmonai is designed by musicians for musicians, targeting hobbyist musicians, sound designers, music producers, researchers, and technologists interested in AI audio generation. The tool is particularly appealing for those interested in experimental music and artistic exploration, as well as anyone wanting to explore AI music generation without budget commitment. The open-source nature promotes collaboration and community engagement among contributors.

The platform runs through Google Colab notebooks, providing a user-friendly interface for accessing the generative audio tools without requiring local installation. Users can export their generated music as audio files and share them. The community-driven approach ensures the tools continue to evolve with contributions from the music and AI community.

Harmonai pricing

Pricing model: Free

Harmonai is entirely free and open-source. All features are included with no paid upgrades available. Users can download the source code, use the tools, and contribute to development without any charge. The platform is browser-based through Google Colab with free access to essential tools. Users may make voluntary donations to support the project. For GPU access, the basic Google Colab plan is free with limited GPU and memory; upgrading to paid Colab plans is optional for more regular experimentation. The Hugging Face organization has a free plan allowing up to 3 spaces and 1 GPU, with paid plans offering more spaces and GPUs.

Harmonai pros

  • Fully open source with no commercial restrictions
  • Completely free to use with no paid tiers or subscriptions
  • Generates original audio from scratch using diffusion models
  • Creates custom infinite sound libraries
  • Six pre-trained models with unique datasets available
  • Supports audio interpolation between two tracks
  • Enables style transfer and audio regeneration
  • Fine-tuning capability with custom datasets
  • Community-driven development with collaborative improvements
  • No reliance on pre-recorded samples or MIDI clips
  • Browser-based through Google Colab, no local installation needed
  • Brings power back to artists而非 corporations
  • Accessible to musicians of all skill levels
  • Copyright-free datasets composed of voluntarily provided music
  • Adjustable parameters for tempo, pitch, volume, and style

Harmonai cons

  • Audio quality is grainy, warbly, and rough compared to commercial services
  • Requires technical knowledge to operate effectively
  • Browser-based only, no desktop or mobile app available
  • Lacks advanced features like multi-track recording or editing
  • Limited community support and documentation
  • Requires Google account and available Google Drive space
  • Basic Google Colab plan has limited GPU and memory
  • Not text-to-audio, has a learning curve to use
  • Model files are several gigabytes each, requiring storage space
  • Only accepts WAV file format for audio uploads

Frequently asked questions about Harmonai

What is Harmonai?

Harmonai is a Stability AI Lab and community-driven organization that releases open-source generative audio tools to make music production more accessible and fun for everyone. It is AI built by musicians for musicians, focused on bringing power back to artists and allowing them to express creativity without limitations.

What is Dance Diffusion?

Dance Diffusion is Harmonai's flagship model, a state-of-the-art AI model that uses machine learning diffusion models to generate high-fidelity and diverse sounds from scratch. It starts from random noise and refines it into music, rather than using pre-existing samples or MIDI clips like commercial tools.

Is Harmonai free to use?

Yes, Harmonai is entirely free and open-source. All features are included with no paid tiers or subscription costs. Users can download the source code, use the tools, and contribute to development without any charge. Voluntary donations are encouraged to support the project.

How do I start using Dance Diffusion?

You need a Google account with available Google Drive space. Access the Dance Diffusion notebook on GitHub which opens in Google Colab. Run the setup by installing dependencies, check GPU status, select a model from the six available options, choose a sampler, and generate new sounds or regenerate your own audio files.

What models are available in Dance Diffusion?

There are six pre-trained models in v0.12: Glitch-440k (glitch and noise samples), Jmann-small-190k and Jmann-large-580k (Jonathan Mann's Song A Day project), Maestro-150k (piano recordings from Google Magenta's MAESTRO dataset), Unlocked-250k (Internet Archive music), and Honk-140k (Canadian Geese wildlife recordings).

Can I fine-tune Dance Diffusion with my own data?

Yes, users can fine-tune the model with custom datasets, allowing them to train it on specific sounds or styles. This is an advanced feature that provides high degrees of creative control for users who want to create personalized audio generation models.

What audio file formats does Harmonai support?

Harmonai currently only accepts WAV files for audio uploads. When uploading audio files to the sample generator in Google Colab, the file path must end with '.wav' extension. Files hosted on Google Drive or other file hosting sites won't work if they don't provide a direct .wav link.

What is the audio quality like?

The audio quality is grainy, warbly, and often sounds a bit weird compared to polished commercial services. This is because Dance Diffusion is still developing and synthesizes audio at a fundamental level from random noise. The quality reflects that it's just getting started, but improves with later iterations.

Can I interpolate between audio files?

Yes, Dance Diffusion supports audio interpolation between two audio files. This feature weaves together separate tracks by combining the style of both tracks, allowing users to hear elements from both audio clips come together in the generated output.

Who is Harmonai for?

Harmonai is ideal for hobbyist musicians and sound designers exploring AI audio generation without budget commitment, music producers interested in experimental music, researchers and technologists studying generative audio models, and anyone who wants to create unique musical compositions without relying on traditional sound libraries.

Categories

Use cases

Browse all AI tools on NeedAnAI