By ShinobiTools Team · Last updated: 1 October 2026
Drop a recording or a video below and Voice only keeps the human voice and pulls the sound around it down. We measured it on barking dogs and birdsong. It works on the sound of a video too, so you can isolate the voice from a video without an editor. No sign-up, no watermark, 10 files a day. One thing it does not do reliably: pick one person out of several people talking at once.
Voice only keeps the human voice and pulls barking and birdsong down. Two people talking at the same time? No setting here reliably separates two voices: Voice only can lose or swap the voice, so try Standard clean-up and listen to the result.
Video (MP4, MOV, MKV, WEBM) or audio (MP3, WAV, M4A, FLAC, OGG), up to 45 minutes and 1 GB. A video comes back as the same video with the cleaned sound; the picture is copied, not re-encoded. You also get the cleaned MP3 and a lossless WAV.
No watermark, no sign-up, 10 files a day. Files up to 45 minutes free. A pass allows longer files and more files a day.
Every 10 dB sounds roughly half as loud. These numbers come from a small set of test recordings where we knew exactly what the clean voice sounded like; they are not reached on every recording.
Upload the MP4, MOV, MKV or WEBM as it is. You get the same video back with the voice kept and the background pulled down, and the picture is copied untouched. The cleaned MP3 and WAV are there as separate downloads if you want to take the voice into an editor.
If "background voices" means a second person talking clearly near the microphone, that is the one case Voice only gets wrong. A second voice is speech too. In our tests with two people talking over each other, the second voice came down by only about 7 dB on average, and on two of four test recordings it lost or swapped the voice. We expect the same with a TV or radio with people talking on it. So we don't promise to remove background voices from a video.
No setting here reliably separates two voices. For a recording with two people talking at once, try Standard clean-up and listen to the result. Need to know who said what? ScribeGrab labels the speakers in a transcript.
| Your recording | Pick |
|---|---|
| One person, with a dog or birds behind them | Voice only |
| Two or more people talking at the same time | Standard clean-up, and listen to the result |
| Only steady hum, hiss or a fan | Standard clean-up |
| Singing over music | neither: use StemGrab's acapella extractor |
If Voice only can't run at that moment, your file gets our other clean-up instead and the result says so. Mostly hum and hiss? The voice enhancer and remove background noise from video pages explain the regular settings.
| Length per file | Files a day | Voice only | |
|---|---|---|---|
| Free, no sign-up | 45 minutes | 10 | yes |
| Free monthly pass (email) | 2 hours | 25 | yes |
| ShinobiTools pass | 3 hours | 50 | yes |
No watermark on any tier. Your upload is deleted right after processing and the result is wiped within 45 minutes.
Drop the file in the tool on this page with Voice only selected and download the result. It is free for 10 files a day, up to 45 minutes each, with no sign-up and no watermark.
Yes. Upload the video and you get the same video back with the voice kept and the background pulled down. The picture is copied as it is, not re-encoded.
Only partly. It pulls down sounds that are not a human voice, such as barking and birdsong. A second person talking clearly is speech too: Voice only may keep that voice, or lose or swap yours.
Not reliably. With two people talking over each other, the second voice came down by only about 7 dB on average in our tests, and on two of four test recordings the voice was lost or swapped. No setting here reliably separates two voices.
It is built to keep the voice. On clean speech in our tests the words stayed the same and the level moved by about 1 dB at most.
Those settings remove noise but are built to keep anything that sounds like a voice, so a bark or birdsong can stay partly audible. Voice only keeps the human voice and pulls the rest down. Standard clean-up is the same automatic clean-up as on our other pages.
No, it is built for speech, not music. To pull the vocals out of a song use StemGrab's acapella extractor.
The upload is deleted right after processing and the cleaned file is wiped within 45 minutes. See privacy.