YAQUIZO
Create free

Remove Silence From Audio

Take the dead air out of a podcast, interview or voice-over — with the threshold measured from your own recording, the pauses shortened rather than deleted, and every join crossfaded so nothing clicks. Nothing is uploaded.

Drop an audio or video file here

MP3, WAV, M4A, FLAC, OGG — or a video, to work on its soundtrack. Nothing is uploaded.

Removing silence sounds like the simplest edit there is: pick a level, delete everything below it, glue the rest together. That is the version most tools ship, and it fails in four fairly specific ways — it clicks at every join, it chops the tails off words, it makes speech sound hurried, and it asks you for a threshold nobody can sensibly guess.

Each of those has a known fix, and this tool applies all four. The joins are crossfaded. Speech opens and closes the gate at different levels, so a fading word cannot flicker it shut. Long pauses are shortened rather than deleted, because pauses are phrasing. And the threshold comes from measuring your recording’s own noise floor instead of asking you to try numbers.

The other half of it is showing the work: the waveform shades every cut before anything happens, the threshold is drawn where it actually falls, and the result plays next to the original. An automatic edit you cannot inspect is one you have to listen to all the way through anyway.

How to use it

  1. 1. Drop in the recording

    Audio or video — a video is handled by working on its soundtrack. It is decoded on your own machine and never uploaded, so an unreleased episode or a client interview stays where it is.

  2. 2. Look at what will be cut

    The waveform shades every section marked for removal and draws the detection threshold across it as a line. You can see whether a quiet word is about to go before anything is changed.

  3. 3. Choose a starting point

    Podcast keeps the phrasing, tight edit is for tutorials and voice-over, jump cut removes every pause, lecture only takes out long dead air. Each one is a set of the controls below, not a black box.

  4. 4. Adjust if it looks wrong

    The shortest pause worth touching, how much of it survives, and the breathing room kept either side of speech. The shading updates as you move them, because detection is fast and rendering is what waits.

  5. 5. Remove and listen back

    The result plays next to the original on the same panel, and downloads on its own. Listening to a join is the only real check.

The four ways this normally goes wrong

It clicks. Cutting a waveform at an arbitrary sample leaves a step, and a few hundred steps is a recording full of ticks. On a test signal, a hard cut leaves a jump about twenty-five times the natural sample-to-sample step; crossfading the join brings that under twice, which is inaudible. The fade uses an equal-power curve rather than a linear one, because the two sides of a join are uncorrelated and a linear fade between uncorrelated signals dips in the middle.

It chops words. With one threshold, the decay at the end of a word crosses it repeatedly and the gate chatters. Here speech has to rise above one level to count as speech and fall below a lower one to count as a pause, which is how a noise gate has worked for decades and for the same reason.

It sounds rushed. Pauses are not waste; they are where sentences end. Speech with every gap removed is recognisably machine-edited. The default shortens long pauses to a set length rather than deleting them, and takes the time out of the middle of the gap so it shrinks symmetrically.

It asks for a threshold. A level that works on one recording is wrong on the next, because noise floors differ by tens of decibels. This measures yours — as a low percentile of the envelope, so a run of digital silence at the head of the file cannot drag it to minus infinity — and sets the threshold a fixed distance above it.

Which preset to start from

Preset Does For
PodcastShortens pauses over 0.35 s to 0.4 sInterviews, conversation. The safe default.
Tight editShortens pauses over 0.25 s to 0.2 sTutorials, voice-over, narration.
Jump cutDeletes every pause over 0.2 sFast-paced video. Obviously edited, deliberately.
LectureOnly touches pauses over 1 sTalks and meetings, where delivery is slower.
Trim endsStart and end onlyA clean take that just needs topping and tailing.
NoisyWider margin above the floorAudible hiss, fan or room tone.

Frequently asked questions

What threshold should I use? +

None — that is the point. Every recording has a different noise floor, so a number that works on one is wrong on the next, and "try −40 dB" is not advice. This measures the noise floor of your actual file and sets the threshold a fixed distance above it. The measured floor and the resulting threshold are both shown, and you can override with a fixed level if you would rather.

Is my audio uploaded? +

No. Everything runs in your browser, so nothing is transmitted and there is no file size limit. That matters more here than for most tools: this is the sort of thing you run on an unreleased episode or a confidential interview.

Why does it shorten pauses instead of deleting them? +

Because speech with no pauses sounds wrong. The pauses carry the phrasing — they are where a sentence ends and a thought turns — and removing all of them produces something that sounds hurried and obviously machine-edited, in a way listeners notice but cannot name. Shortening a two-second gap to half a second takes out the dead air and keeps the rhythm. Deleting them entirely is available as the jump-cut preset, which is the right choice for fast-paced video.

Will the cuts click? +

No, because every join is crossfaded. This is the single commonest failure of automatic silence removers: cutting a waveform at an arbitrary sample leaves a step discontinuity, and a few hundred of those is a recording full of ticks. Measured on a test signal, a hard cut leaves a jump about twenty-five times the natural sample-to-sample step; with the crossfade, it drops to well under twice — inaudible. The crossfade length is adjustable, and setting it to zero is measurably worse.

Will it cut off the ends of words? +

It is designed not to, in two ways. Speech opens and closes the gate at different levels, so the fading tail of a word cannot flicker it shut mid-syllable. And a configurable amount of silence is kept either side of every retained section, so attacks are not clipped. If you set the breathing room to zero and the threshold high, it will absolutely clip words — which is why the waveform shows you the cuts before you commit to them.

What if the recording is noisy? +

Then silence detection gets harder, because there is less difference between someone speaking and the room. The tool measures that gap and warns you when there is under about 12 dB between the noise floor and the loudest moment — below that, any threshold is a guess. The fix is to remove the noise first and then run this, which is a much cleaner result than any threshold can achieve on its own.

Can it just trim the start and end? +

Yes — there is a preset for exactly that. Cutting the silence before you started talking and after you stopped, without touching anything in between, is one of the most common reasons people look for this.

Does it work on video? +

It works on a video file’s soundtrack and gives you back audio. It does not currently cut the picture to match, so for video editing you would use the result as a guide rather than a finished track.

How much time will it save? +

That depends entirely on the recording, which is why the tool tells you rather than promising a figure. A tightly delivered read might lose five per cent; an unrehearsed interview with long thinking pauses can lose twenty or more. If it reports removing over half the file, it says so as a warning rather than a triumph — that usually means the voice is quiet relative to the background and some speech may have gone with the silence.

What formats can I save as? +

M4A for something small that plays everywhere, WAV for the exact samples with no further compression, or Opus. M4A at 128 kbps is the sensible default for spoken word.

Does it change the volume or the sound? +

No. Nothing is compressed, levelled or filtered — sections are removed and the joins are crossfaded, and everything that remains is bit-for-bit what you gave it, apart from the few milliseconds inside each fade. If you also want the levels set, run the audio editor afterwards.

Other free tools