Audio DeSilencer is a small Python package and command-line tool for shortening long silent pauses without processing the audio that is kept.
The core rule is simple:
If a source interval is retained, Audio DeSilencer does not normalize, limit, EQ, compress, de-ess, crossfade, or otherwise modify it. It only removes selected time ranges.
The tool is designed for speech recordings, meetings, voice notes, podcasts, screen recordings, interviews, lectures, and other audio where long pauses should be shortened while preserving the original signal.
Audio DeSilencer:
- detects continuous quiet regions
- keeps a small amount of natural pause around speech
- removes only the middle of accepted silent regions
- leaves retained audio untouched
- writes a cleaned
*_non_silent.wav - optionally writes the removed material as
*_silent.wavfor inspection - writes source-time timeline files showing exactly what was removed and retained
- supports fixed or automatic threshold detection
- handles mono, stereo, and multichannel audio conservatively
- uses FFmpeg for decoding, detection, trimming, and output
It deliberately does not perform mastering or restoration. There is no normalization, limiter, denoiser, EQ, compressor, de-esser, voice enhancement, or automatic loudness processing.
- Python 3.10 or newer
- Git
- FFmpeg and FFprobe available on
PATH
No third-party Python runtime package is required by Audio DeSilencer itself.
winget install Gyan.FFmpegClose and reopen PowerShell after installation if necessary, then verify:
ffmpeg -version
ffprobe -versionbrew install ffmpegVerify:
ffmpeg -version
ffprobe -versionsudo apt update
sudo apt install ffmpegVerify:
ffmpeg -version
ffprobe -versionThe recommended installation method is to install directly from a checked-out copy of this repository.
git clone https://github.com/BTawaifi/Audio-DeSilencer.git
cd Audio-DeSilencer
py -m pip install --upgrade pip
py -m pip install .Verify the installed command:
audio-desilencer --help
where.exe audio-desilencergit clone https://github.com/BTawaifi/Audio-DeSilencer.git
cd Audio-DeSilencer
python3 -m pip install --upgrade pip
python3 -m pip install .Verify:
audio-desilencer --help
which audio-desilencerIf you are modifying the source code, install it in editable mode:
py -m pip install -e .python3 -m pip install -e .With an editable install, changes in the repository are used without rebuilding and reinstalling the package each time.
If you already cloned the repository:
cd D:\path\to\Audio-DeSilencer
git pull
py -m pip install --force-reinstall .cd /path/to/Audio-DeSilencer
git pull
python3 -m pip install --force-reinstall .If the CLI still looks old after reinstalling, see Troubleshooting.
For normal speech recordings, the CLI has useful zero-configuration defaults.
audio-desilencer "C:\Users\User\Documents\Sound recordings\Recording (2).m4a"or on macOS/Linux:
audio-desilencer "recording.m4a"The default CLI behavior is:
threshold mode: auto
minimum continuous silence: 700 ms
silence retained per pause: 150 ms
auto hysteresis: 3 dB
output format: WAV
output folder: output
In most voice-recording cases, you should start with no tuning flags at all.
For an input named:
recording.m4a
the default output folder contains:
output/
├── recording_non_silent.wav
├── recording_silent.wav
├── recording_non_silent_parts.txt
└── recording_silent_parts.txt
This is the main result: the original recording with accepted long pauses shortened.
It still contains the small amount of pause requested by --target_silence_len so speech does not get slammed directly together.
This is an inspection artifact containing the source ranges that were removed, concatenated together.
It is not intended to sound like a normal continuous recording. Its purpose is to let you inspect what the tool removed.
Contains the source-time ranges actually removed.
Example:
[(9824, 10599), (19576, 20330)]
All values are milliseconds in the original source timeline.
Contains the source-time ranges retained in the cleaned output.
audio-desilencer INPUT_FILE [options]
Show all options:
audio-desilencer --helpRequired path to the source audio file.
Examples:
audio-desilencer "voice.m4a"
audio-desilencer "C:\Recordings\meeting.ogg"FFmpeg determines which formats are supported by your local installation. Common formats such as WAV, M4A/AAC, MP3, OGG/Opus, FLAC, and many others are generally supported by standard FFmpeg builds.
Folder where generated files are written.
Default:
output
Example:
audio-desilencer "voice.m4a" --output_folder "D:\Processed Audio"Controls what level is considered quiet.
CLI default:
auto
audio-desilencer input.m4a --threshold autoAuto mode is deterministic signal analysis. It does not use AI, a trained model, cloud service, internet connection, or external dataset.
It examines short peak-level windows and estimates a conservative threshold from the quieter part of the recording.
Important safety behavior:
- a loud sample makes the window active
- activity on any channel protects the interval
- auto mode does not choose a threshold more aggressive than the built-in
-38 dBFSsafety ceiling - auto mode uses exit hysteresis by default
Use a numeric dBFS value when you want explicit, reproducible detection:
audio-desilencer input.m4a --threshold -38A more negative value is more conservative and classifies less material as silence:
-42 dBFS -> safer for quiet speech
-38 dBFS -> practical fixed speech starting point
-32 dBFS -> more aggressive; use carefully
Minimum continuous quiet duration required before a region can be considered silence.
Default:
700 ms
Example:
audio-desilencer input.m4a --min_silence_len 1000Higher values remove only longer pauses. Lower values make the tool more aggressive.
Recommended speech range:
500-1000 ms
How much total pause remains after an accepted long silence is shortened.
Default:
150 ms
Examples:
# Tighter pacing
audio-desilencer input.m4a --target_silence_len 50
# Default natural speech pacing
audio-desilencer input.m4a --target_silence_len 150
# Preserve more breathing room
audio-desilencer input.m4a --target_silence_len 250
# Remove accepted silence completely
audio-desilencer input.m4a --target_silence_len 0For normal speech, 100-200 ms is usually a good range.
Optional difference between entering and leaving the silence state.
Auto mode uses 3 dB by default.
For a fixed threshold:
audio-desilencer input.m4a --threshold -38 --hysteresis_db 3Hysteresis helps prevent rapid switching when the signal hovers near the threshold.
If you use a numeric threshold without --hysteresis_db, the tool keeps the direct FFmpeg fixed-threshold detector path.
Output audio format.
Default:
wav
Example:
audio-desilencer input.m4a --output_format mp3WAV is recommended when quality and validation matter because the tool writes 32-bit floating-point PCM. Lossy formats such as MP3 must be re-encoded and therefore cannot be sample-identical to the decoded source.
Override the base filename used for generated outputs.
Example:
audio-desilencer interview.m4a --output_stem cleaned_interviewProduces names such as:
cleaned_interview_non_silent.wav
cleaned_interview_silent.wav
Use the defaults:
audio-desilencer input.m4aEquivalent conceptually to:
threshold: auto
min silence: 700 ms
target pause: 150 ms
auto hysteresis: 3 dB
Prefer auto first. If you need a fixed threshold, use a more negative value:
audio-desilencer input.m4a --threshold -42First reduce the retained pause:
audio-desilencer input.m4a --target_silence_len 80If long pauses are still not detected, lower the required duration:
audio-desilencer input.m4a --min_silence_len 500With a fixed threshold, you can cautiously make the threshold less negative:
audio-desilencer input.m4a --threshold -35Do this carefully because aggressive thresholds can classify quiet speech as silence.
Make detection more conservative:
audio-desilencer input.m4a --threshold -42 --min_silence_len 800Or return to auto mode:
audio-desilencer input.m4a --threshold autoYou can also preserve more pause:
audio-desilencer input.m4a --target_silence_len 250A conservative example:
audio-desilencer meeting.m4a \
--threshold auto \
--min_silence_len 900 \
--target_silence_len 180PowerShell equivalent:
audio-desilencer "meeting.m4a" `
--threshold auto `
--min_silence_len 900 `
--target_silence_len 180Audio DeSilencer has two deterministic detection paths.
Auto mode:
- decodes the audio through FFmpeg into floating-point samples for analysis
- measures short peak-level windows
- estimates the low-level/noise portion of the recording
- chooses a conservative silence threshold
- uses hysteresis so the detector does not rapidly switch state around the threshold
- requires continuous quiet for at least
--min_silence_len
The detector is intentionally conservative. If uncertain, it should leave silence rather than remove possible speech.
A numeric threshold without hysteresis uses FFmpeg's silencedetect path directly.
Example:
audio-desilencer input.m4a --threshold -38This is useful when you need repeatable behavior with an explicitly chosen threshold.
The tool does not reconstruct speech from arbitrary fragments.
Instead:
source recording
|
v
detect confirmed long quiet region
|
v
keep a small pause next to speech
|
v
remove only the middle of that quiet region
|
v
concatenate retained source-time spans
For example, if a detected pause is 1.5 seconds and the target pause is 150 ms, the tool removes the middle portion and leaves approximately 75 ms on each side of the cut.
This keeps edits inside quiet material rather than directly on speech boundaries.
Audio DeSilencer treats silence removal as an editing problem, not a mastering problem.
The rendering graph uses timing operations only:
atrim -> timestamp reset -> concat
Retained audio is not intentionally subjected to:
- normalization
- limiting
- gain changes
- compression
- EQ
- high-pass filtering
- de-essing
- denoising
- crossfades
- voice enhancement
For WAV output, 32-bit floating-point PCM is used so over-range floating-point decoded samples are not unnecessarily forced through an integer PCM ceiling.
Lossy output formats still require codec re-encoding, so use WAV when checking waveform integrity or when you want a lossless intermediate.
Detection is conservative across channels.
A window is considered quiet only when all channels are quiet. If one channel contains activity, the interval is protected.
This avoids deleting useful material from stereo recordings where speech or another important signal exists primarily on one side.
Very long recordings or recordings with many pauses can create large FFmpeg filter graphs.
Audio DeSilencer automatically switches large graphs to a temporary FFmpeg filter-script file instead of putting the entire graph on the command line. This also avoids operating-system command-length problems, especially on Windows.
The temporary filter script is deleted after rendering.
from audio_desilencer import AudioProcessor
processor = AudioProcessor("input.m4a")
result = processor.process_audio(
threshold="auto",
min_silence_len=700,
target_silence_len=150,
output_folder="output",
output_format="wav",
)
print(result.non_silent_audio_path)
print(result.silent_audio_path)
print(result.detected_silence_ranges)
print(result.silent_ranges)
print(result.non_silent_ranges)
print(result.resolved_threshold_dbfs)
print(result.detector)The CLI defaults to threshold="auto" for convenient speech processing.
The lower-level Python AudioProcessor.process_audio() API keeps its fixed numeric threshold default for compatibility with existing callers. If you want the same adaptive behavior as the CLI, pass:
threshold="auto"silent_ranges = processor.detect_silence_ranges(
min_silence_len=700,
threshold="auto",
)silent_ranges = processor.detect_silence_ranges(
min_silence_len=700,
threshold=-38,
)The compatibility API remains available:
non_silent_ranges = processor.split_audio_by_silence(
min_silence_len=700,
threshold=-38,
)ProcessingResult exposes:
silent_audio_path
non_silent_audio_path
silent_timeline_path
non_silent_timeline_path
silent_ranges
non_silent_ranges
detected_silence_ranges
resolved_threshold_dbfs
detector
silent_ranges means ranges actually removed, not every region that was initially detected as low-energy.
Use these steps when you want a wheel/source distribution instead of installing directly from the checkout.
From the repository root:
py -m pip install --upgrade build
Remove-Item -Recurse -Force dist -ErrorAction SilentlyContinue
py -m buildGenerated files appear in:
dist/
Typically:
audio_desilencer-1.1.0-py3-none-any.whl
audio_desilencer-1.1.0.tar.gz
Install the locally built wheel:
py -m pip install --force-reinstall .\dist\audio_desilencer-1.1.0-py3-none-any.whlOr when only one wheel exists:
py -m pip install --force-reinstall .\dist\*.whlpython3 -m pip install --upgrade build
rm -rf dist
python3 -m build
python3 -m pip install --force-reinstall dist/*.whlaudio-desilencer --helpWindows:
py -m pip show Audio-DeSilencer
where.exe audio-desilencermacOS/Linux:
python3 -m pip show Audio-DeSilencer
which audio-desilencerFFmpeg and FFprobe must be available before running tests.
From the repository root:
python -m unittest discover -s tests -vOn Windows you can use:
py -m unittest discover -s tests -vThe regression suite covers areas including:
- leading, internal, and trailing silence
- fully silent and fully active files
- middle-of-pause removal
- automatic threshold estimation
- fixed threshold detection
- hysteresis behavior
- multichannel safety
- retained decoded PCM preservation
- floating-point over-range preservation
- large edit graphs
- output naming and path containment
- timeline semantics
- CLI success and error handling
GitHub Actions runs tests across supported Python versions and separately validates the built package.
A typical local development setup:
git clone https://github.com/BTawaifi/Audio-DeSilencer.git
cd Audio-DeSilencer
py -m pip install -e .
py -m unittest discover -s tests -vAfter making code changes, rerun the test suite before building a wheel.
To inspect the CLI from the current editable checkout:
audio-desilencer --helpIf you see an error similar to:
argument --threshold: invalid int value: 'auto'
your shell is running an older installed CLI.
Update the repository and reinstall from source:
git pull
py -m pip uninstall Audio-DeSilencer -y
py -m pip install --force-reinstall .Then inspect which executable is being used:
where.exe audio-desilencerVerify the help now includes:
--threshold
--hysteresis_db
--target_silence_len
First verify the package is installed:
py -m pip show Audio-DeSilencerThen inspect your Python Scripts directory and PATH.
On Windows, where.exe audio-desilencer is useful once the command is discoverable.
Verify:
ffmpeg -version
ffprobe -versionIf those commands fail, install FFmpeg and reopen your shell so the updated PATH is loaded.
Use a more conservative fixed threshold and/or longer minimum silence:
audio-desilencer input.m4a --threshold -42 --min_silence_len 900You can also keep more natural pause:
audio-desilencer input.m4a --target_silence_len 250Try a shorter target pause first:
audio-desilencer input.m4a --target_silence_len 80Then, if needed, reduce minimum silence duration:
audio-desilencer input.m4a --min_silence_len 500For fixed detection, cautiously use a less-negative threshold:
audio-desilencer input.m4a --threshold -35MP3 is lossy and must be encoded again after editing.
Use WAV when validating signal preservation:
audio-desilencer input.m4a --output_format wavThat file contains removed ranges concatenated together. It is only an inspection artifact.
Listen to *_non_silent.wav to judge the cleaned recording.
The tool automatically uses an FFmpeg filter-script file when the edit graph is large, so normal use should not require any manual workaround.
CLI / Python caller
|
v
AudioProcessor
├── ffprobe source metadata
├── deterministic auto analysis
│ OR fixed FFmpeg silence detection
├── normalize detected silence ranges
├── shorten only accepted silence interiors
├── compute retained source ranges
├── FFmpeg trim + concatenate retained spans
├── export removed spans for inspection
└── write source-time timelines
The detector decides where editing is safe. The renderer does not try to hide detector mistakes by processing the voice.
- Detection is deterministic signal analysis, not semantic speech recognition.
- No ML/VAD model, training data, cloud API, or internet service is required.
- Auto mode is deliberately conservative and may leave extra silence rather than risk deleting quiet speech.
- Extremely noisy recordings may still need a manually chosen threshold.
- The tool does not repair clipping already present in the source.
- The tool does not normalize loudness.
- Codec support depends on the FFmpeg build installed on the host.
- Lossy output formats cannot be sample-identical to the source because they must be re-encoded.
py -m pip uninstall Audio-DeSilencerpython3 -m pip uninstall Audio-DeSilencerMIT