Audioshake

Audioshake is an AI tool designed to help musicians, labels, publishers, and other stakeholders unlock new potential in audio recordings by separating tracks into individual elements.

Last verified:

Visit Audioshake

What is Audioshake?

AudioShake is an AI-powered audio separation platform that uses deep learning to recognize and isolate different components within audio files, such as vocals, drums, bass, instruments, dialogue, music, and sound effects. The technology transforms mixed audio into editable stems without generating new content, enabling users to split fully mixed recordings into individual elements for remixing, mastering, localization, and other professional workflows.

Key features include instrument stem separation (vocals, drums, bass, guitar, piano, winds, strings), dialogue/music/effects separation for film and TV, multi-speaker separation for overlapping voices in podcasts, and lyric transcription with word-by-word time stamping for karaoke experiences. The platform supports audio up to 192kHz and exports in WAV, MP3, AAC, FLAC, AIFF, and PCM formats, with transcriptions available as JSON or TXT files.

AudioShake serves music labels, publishers, producers, DJs, engineers, film/TV studios, broadcasters, dubbing studios, developers, indie artists, and content creators. Major customers include Disney Music Group, Universal, Sony, Warner, Paramount, BET, NFL Films, Netflix, and companies like cielo24, OOONA, and Yella Umbrella. The platform won Sony's Demixing Challenge, outperforming 40 other teams including big tech companies and research institutes.

The tool is used for mixing and mastering (creating Dolby Atmos and Sony 360 immersive mixes, remastering classic recordings like Jackson 5 and Whitney Houston), localization and captioning (improving transcription accuracy by 25%+), interactive audio for gaming and fitness, sync licensing (creating instrumentals for Disney/Netflix trailers), copyright compliance (sample clearance, royalty allocation), and A/V editing (removing copyrighted music, cleaning broadcasts).

Audioshake pricing

Pricing model: Free

AudioShake offers usage-based pricing with one-off processing around $1 per minute and API usage in the cents per minute. Discounts are available for bulk processing. A free tier is available: new accounts receive 10 free credits to start building immediately. AudioShake Indie is designed for independent artists and small labels. AudioShake Live is an on-demand platform for industry professionals (labels, publishers). Enterprise API/SDK integrations are available for large companies. Free plan exists for small projects to try the service. Request a demo for custom enterprise pricing.

Audioshake pros

  • Industry-leading stem separation quality that won Sony's Demixing Challenge
  • Separates vocals, drums, bass, guitar, piano, winds, strings, and more
  • Supports audio up to 192kHz with professional broadcast quality
  • Multiple export formats: WAV, MP3, AAC, FLAC, AIFF, PCM, JSON, TXT
  • Improves transcription accuracy by 25% or more for dialogue cleaning
  • Trusted by major labels: Universal, Sony, Warner, Disney Music Group
  • API and SDK available for iOS, MacOS, Windows, Android, Linux
  • Multi-speaker separation isolates overlapping voices in podcasts
  • Word-by-word lyric alignment for karaoke experiences
  • Dialogue, music, and effects separation for film/TV dubbing workflows
  • Creates instrumentals for sync licensing in seconds
  • Does not generate new content - only separates existing audio
  • On-device separation available for edge deployment
  • Processes millions of hours of audio annually at scale
  • Drag-and-drop platform (AudioShake Live) for labels and publishers
  • AudioShake Indie plan for independent artists and small labels
  • Preserves integrity of original performance in stems

Audioshake cons

  • Usage-based pricing around $1 per minute for one-off processing
  • API usage costs in cents per minute can add up for large volumes
  • No speech transcription - only lyric transcription available
  • Output download links expire after one hour
  • Tasks process asynchronously requiring polling for completion
  • API key cannot be viewed again after creation
  • Indie plan may have limitations compared to Live for enterprise
  • Requires account creation and API key setup for developers

Frequently asked questions about Audioshake

How does AudioShake's technology work?

AudioShake uses AI to recognize different components in a piece of audio, such as drums in a rock song or dialogue in film. The technology then isolates that stem so you can use it for new purposes including sampling, synch licensing, re-mastering, re-mixes, localization, and more. It is source separation technology that takes existing audio and separates it into different components without generating new content.

What kinds of sound separation do you offer?

AudioShake offers multiple models: Music (separate different instruments or create an instrumental, separate multiple singers), Dialogue/Music/Effects (separate speech/dialogue from background audio or effects from music), Multi-speaker Separation (separate overlapping speakers in podcasts/videos/speech files), and Lyric Transcription & Alignment (transcribe lyrics with word-by-word time stamping for karaoke experiences).

Is AudioShake available via API or on-device?

Yes, all separations are available via API, and many are also available on-device. The documentation site provides integration details. SDKs are available for iOS, MacOS, Windows, Android, and Linux platforms for edge device deployment.

What file formats and resolutions do you support?

AudioShake matches file inputs and supports up to 192kHz resolution. Export formats include WAV, MP3, AAC, FLAC, AIFF, and PCM for audio. For transcriptions, text can be exported as JSON or TXT files.

How can I get my audio stemmed?

For music professionals, AudioShake Live is an on-demand platform where you can upload songs and create stems. Get in touch for a demo and free trial. For independent artists and labels, AudioShake Indie is available. Music supervisors can use services on Chordal. Dubbing freelancers and studios find the technology embedded in OOONA and Yella Umbrella workflows, and through services including Dubverse and cielo24. Developers can integrate via the API.

What makes AudioShake different from similar tech?

Recent AI advances enabled a big leap in quality. AudioShake's technology is best in the industry and significantly outperforms other offerings. They won Sony's Demixing Challenge, pitting their stem separation against 40 other teams including Big Tech companies, startups, and research institutes. The technology delivers broadcast quality usable for commercial purposes.

Do you do speech transcription and alignment?

No, in transcription AudioShake only focuses on lyrics, which is a different research task from speech. However, they work with many speech transcription and captioning services to clean dialogue before it goes through automated speech recognition (ASR) by extracting clean dialogue stems.

What is AudioShake used for in mixing and mastering?

AudioShake splits music into instrument stems for mixing, mastering, or creating immersive mixes. The stems are used to master live recordings or mix tracks missing stems. They create Dolby Atmos and Sony 360 immersive mixes and have been used for Jackson 5, Nina Simone, Whitney Houston, and more catalog remastering projects.

How does AudioShake help with localization and captioning?

AudioShake cleans dialogue while retaining original music to improve dubbing workflows. Background noise and music interfere with ASR transcription, but AudioShake extracts clean dialogue stems with customer-reported increases of 25% or more in transcription accuracy rates for transcription, captioning, and automatic dubbing, while allowing original music and effects to be reused in localized output.

What customers use AudioShake?

Industry leaders include Disney Music Group, Universal, Sony, Warner Music Group, Paramount, BET, NFL Films, Netflix, EMPIRE, Peermusic, and tech companies. Endorsements come from David Abdo (SVP Disney Music Group), Tony Abrahams (CEO AI Media), Ghazi Shami (CEO EMPIRE), Grammy winners Billy Mann and Rich Keller, and companies like cielo24, Audiosocket, Immersive Mixers, and The Teenage Diplomat.

Categories

Use cases

Browse all AI tools on NeedAnAI