Vocal Remover (Karaoke Maker)
Split a song into vocals and backing track — an AI model runs inside your browser, no upload
1. Pick a song
Splits a song into two tracks — vocals and backing (MR). A source-separation AI model (MDX-Net family) runs directly inside your browser, so the file is never uploaded; only the model (66MB) is downloaded once on first run. Requires a desktop browser with WebGPU support.
What you can do
- ✓Separates vocals and backing track (MR)
- ✓AI model runs in your browser (no upload)
- ✓Preview each track, save as WAV/MP3
- ✓Works with mono sources too
How to use
- 1Add a song file
- 2Click 'Split vocals & backing' and wait
- 3Listen to the stems and download them
Supported formats · MP3 · M4A · WAV · OGG · FLAC → WAV · MP3 (backing/vocals)
Split a song into two tracks — vocals and backing (MR). A widely proven source-separation AI model (MDX-Net family) runs directly inside your browser, so the file is never uploaded. Download the backing track for karaoke and practice, or the vocals for covers and remix practice. Requires a desktop browser with GPU acceleration (WebGPU).
How to use it
- 1Select or drag in a song (MP3, M4A, WAV, OGG, FLAC, up to 12 minutes).
- 2Click 'Split vocals & backing'. The first run downloads the AI model (66MB) once.
- 3Listen to the backing track and vocals, then download each as WAV or MP3.
How it works
The song is cut into ~6-second segments and converted to a spectrogram (a frequency map), where an AI model trained specifically for source separation extracts the backing track. The vocal track is obtained by subtracting the backing from the original.
Everything, including model inference, runs inside your browser on your device's GPU — the file never leaves your device. Only the model file (66MB) is downloaded once and reused.
Time and requirements
Expect tens of seconds to a few minutes for a typical 3-4 minute song, depending on your hardware. You'll need an up-to-date Chrome or Edge on a computer with WebGPU support; older browsers and some phones won't work.
You can use other tabs while it runs — plugging a laptop into power makes it finish faster.
Honest notes on quality
On typical studio mixes you'll get a backing track where the vocal is barely audible. AI separation isn't magic though: stacked harmonies, live recordings, and low-quality sources can leave traces or reverb behind.
If the result disappoints, trying a better-quality source file usually helps the most.
When to use it
Four common situations
Karaoke practice
Make a backing track of a favorite song and sing along. You can also listen to the isolated vocal to study melody and phrasing.
Recording covers
Layer your voice over the separated backing track for a clean cover. Record with Voice Recorder and combine with Audio Merge.
Events and parties
Need an accompaniment fast — a congratulatory song, a talent show? Make one in the browser with nothing to install.
Instrument and ear practice
Instruments stand out on the backing track and the melody line on the vocal track — great for learning parts by ear.
Frequently asked questions
Does it remove vocals completely?+
On typical studio mixes the vocal becomes barely audible. Stacked harmonies, live recordings, and low-quality sources can leave traces.
How is AI separation different from the phase-inversion (L-R) method?+
Many free vocal removers subtract the right channel from the left (L-R) to cancel whatever sits in the center. It's instant, but it has a built-in flaw: center-mixed bass, kick, and lead instruments vanish too, and vocal reverb stays. This tool instead uses an AI model trained for source separation that identifies the vocal in the spectrum and lifts it out, so the instruments survive. The trade-off is a one-time model download and some processing time.
Won't the bass and drums disappear too?+
No — that's the chronic problem of the L-R method. With AI separation, instruments stay in the backing track even when they're mixed dead-center alongside the vocal.
Won't the backing track sound quieter than the original?+
Removing the loudest element of a mix (the vocal) naturally lowers perceived loudness. That's why the backing track's loudness is automatically matched to the original's average level, within a clipping-safe limit.
Is my file uploaded to a server?+
No. The AI model runs inside your browser on your device's GPU, so the file never leaves your device. Only the model file (66MB) is downloaded once.
Why does it need WebGPU?+
The model is computationally heavy — without GPU acceleration it would take far too long. Use an up-to-date Chrome or Edge on a computer.
How long does it take?+
Depends on your hardware — tens of seconds for a 3-4 minute song on a recent PC, up to a few minutes on older machines.
Does mono audio work?+
Yes. Phase-inversion (L-R) tools need distinct stereo channels and go silent on mono, but AI separation analyzes the sound itself, so mono sources work fine.
Can I save just the vocals?+
Yes. You get two tracks — backing (MR) and vocals — each downloadable as WAV or MP3.
Can I publish the karaoke track I made?+
Use it for personal practice and enjoyment. Separation doesn't change the song's copyright, so distributing or publishing the stems requires the rights holder's permission.
Is it free? Any limits?+
Free, no account needed. Because processing happens in browser memory, files are capped at 50MB and 12 minutes — split longer audio with Audio Cut.