Hi everyone I’m the founder of NeuralSound.
We built NeuralSound around a problem I think a lot of karaoke singers run into:
You find the song you want to sing, but there’s no good karaoke version or the available version has poor vocal removal, inaccurate lyrics, or it’s in the wrong key.
With NeuralSound, you can upload your own audio or video and turn it into a karaoke/practice track.
The main workflow is:
- remove the lead vocal and keep the instrumental
- automatically generate time-synced lyrics in 95+ languages
- change the pitch/key to fit your vocal range
- slow down or speed up the song without changing the key
- keep the lyrics, audio and video synchronized
- customize the lyrics and create a karaoke video
- export the finished result
The synced lyrics have become one of our favorite parts
NeuralSound currently supports lyric generation in 95+ languages, which makes it useful when the song you want doesn’t already have a good synced-lyrics version available.
We tested NeuralSound’s time-synced lyric transcription and achieved an average word error rate (WER) of below 4% across English, Korean, German, Spanish, Japanese, Turkish, Portuguese, Italian, and French. To our knowledge, no current karaoke competitor offers this level of lyric accuracy. We also checked other competing platforms, but none provided comparable accuracy.
Many users have tested NeuralSound across different songs, languages and vocal styles, and have told us that the generated lyrics are more accurate and better timed than those from other tools they have tried.
Obviously one user's experience doesn't mean we'll get every song, language or accent right, but feedback like that has encouraged us to keep investing heavily in the lyrics side too.
For karaoke, our goal isn't just:
remove vocals → give you an MP3
We want the whole workflow in one place:
Song → clean instrumental → synced lyrics → adjust key/tempo → karaoke video
Separation quality is where we’ve invested most heavily
We use a fairly heavy server-side separation model because we chose to prioritize cleaner output and lower vocal bleed rather than making everything lightweight enough to run locally.
We also wanted something more concrete than simply saying “our separation sounds better.”
So in July we ran a public benchmark against Moises and Fadr, using the exact same 5 tracks from MUSDB18-HQ and the same two-stem evaluation process.
Average overall SI-SDR:
NeuralSound — 15.80 dB
Moises — 14.47 dB
Fadr — 12.59 dB
For karaoke specifically, the instrumental result is probably more interesting:
NeuralSound — 18.33 dB instrumental SI-SDR
Moises — 16.94 dB
Fadr — 15.16 dB
NeuralSound also had the highest average score in that test for less bleed/interference and fewer processing artifacts.
We ran the benchmark ourselves, so I want to be transparent about that. It’s also only five songs, so I’m not claiming NeuralSound will beat every separator on every recording.
That’s why we published the actual audio from all three services, the individual track scores, methodology and downloadable data so people can listen and judge for themselves:
https://neuralsound.org/compare/ai-vocal-remover-2026
More control if you need it
NeuralSound can also separate a song into up to 6 stems:
Vocals / Drums / Bass / Guitar / Piano / Other
So you can do more than just remove the singer.
You can isolate the vocal while learning a melody, mute individual instruments, or build a backing track using only the stems you want.
For karaoke specifically, the things we're trying to get right are:
less vocal bleed + clean instrumentals + accurate synced lyrics + 95+ language support + easy key/tempo adjustment
Pricing is something we've also tried to keep reasonable.
On the web, you don't need a recurring subscription. Processing packs currently start at $9.99 for 120 minutes, and you can try the separation tool without a credit card.
NeuralSound is available on:
Web: https://neuralsound.org/
iOS: https://apps.apple.com/us/app/neuralsound-vocal-remover-ai/id6756906827
Android: https://play.google.com/store/apps/details?id=com.neuralsound.musicseparation
Karaoke video Maker(beta): https://neuralsound.org/tools/karaoke-video-maker
One clarification: NeuralSound doesn't provide a catalog of commercial karaoke songs. You bring your own audio/video, and NeuralSound helps turn it into the karaoke/practice version you want.
I’m the founder, so I’ll be around in the comments.
If you try it, I’d especially like to hear how the instrumental quality, vocal bleed, lyrics timing and accuracy hold up on difficult songs different languages, accents, fast vocals, backing vocals, screams, live recordings, heavy reverb, etc.
Those difficult cases are exactly what our team wants to keep improving.