DeepFilterNet3 is the free, local noise remover to use on recorded voice: it lifted speech quality on every noise we tested, from keyboard clatter to mains hum, and it cleaned 10 minutes of audio in 27 seconds on one CPU core. The catch is that at its default setting it also deletes a music bed almost completely. We ran it against RNNoise (the engine behind OBS's noise filter, in two versions) and the two FFmpeg filters, afftdn and arnndn, on 824 clips from a standard speech benchmark and 800 clips of our own test mixes, all on a CPU with no GPU.
Every result below comes from files we processed and scored ourselves on 30 September 2026. Speech quality is PESQ, the ITU-T P.862.2 score from 1 (bad) to 4.64 (identical to the clean recording). The scripts and raw results are listed at the end.
Quick Picks
- Pick DeepFilterNet3 if you clean up recorded voice after the fact: podcasts, voiceovers, interviews, screen recordings. At full strength it scored highest on eight of our nine tests, and its capped setting took the ninth, a clean recording with no noise at all.
- Pick DeepFilterNet3 capped at 12 dB (
-a 12) if there is music, ambience or room tone under the voice that you want to keep. It removes exactly 12 dB of the bed instead of all of it, and it scored below the untouched noisy clip on 0 of 824 benchmark clips. - Pick RNNoise 0.2 if you need live suppression on a very light CPU budget, and switch it off when the room is quiet. It beat the older RNNoise model that OBS ships on 577 of 824 clips, but both versions pulled clean speech from 4.64 down to about 3.9.
Skip FFmpeg's afftdn at its defaults. It changed almost nothing on the noises that bother creators most, and even run on 10 minutes of continuous audio with a hand-set noise floor it gained a fraction of what the neural models did. And skip all of them before transcription: none made Whisper more accurate, and RNNoise 0.2 raised its word error rate from 3.75% to 5.68%.
Detailed Comparison
Four tools, two RNNoise models, seven settings. afftdn is FFmpeg's classic spectral filter, tested at its defaults and with noise tracking on (tn=1); the tracking version scored within 0.01 of the defaults on every test, so the table shows the defaults. arnndn is FFmpeg's built-in RNNoise-style filter, which needs a model file; we used "somnolent-hogwash" (sh.rnnn), the rnnoise-models entry for speech in recording noise. RNNoise ran as the original model OBS ships and as the v0.2 release from April 2024. DeepFilterNet3 ran through the deep-filter 0.5.6 command-line binary, at full strength and capped at 12 dB.
| Test (mean PESQ) | Untouched | afftdn | arnndn | RNNoise (OBS) | RNNoise 0.2 | DeepFilterNet3 | DFN3, 12 dB cap |
|---|---|---|---|---|---|---|---|
| Benchmark, 824 clips | 1.97 | 2.09 | 2.15 | 2.32 | 2.46 | 3.05 | 2.67 |
| Keyboard, 5 dB | 1.21 | 1.21 | 1.34 | 1.42 | 1.55 | 2.67 | 1.48 |
| Keyboard, 15 dB | 1.45 | 1.48 | 1.61 | 1.82 | 1.99 | 3.19 | 2.08 |
| Dog bark, 5 dB | 1.73 | 1.71 | 1.84 | 1.71 | 1.83 | 3.09 | 2.18 |
| Dog bark, 15 dB | 2.20 | 2.21 | 2.33 | 2.11 | 2.27 | 3.57 | 2.87 |
| 60 Hz hum, 10 dB | 1.62 | 1.58 | 2.08 | 2.27 | 2.26 | 3.30 | 2.58 |
| Mic hiss, 10 dB | 1.23 | 1.25 | 1.86 | 1.97 | 1.98 | 2.51 | 1.85 |
| Clean voice, no noise | 4.64 | 4.45 | 4.39 | 3.95 | 3.91 | 4.48 | 4.52 |
| Voice over a music bed | 1.61 | 1.59 | 1.83 | 1.84 | 1.92 | 3.12 | 2.32 |
The dB figure is the signal-to-noise ratio: at 5 dB the noise is loud, at 15 dB it is in the background. Every creator test used the same 100 clean sentences, so each row compares the tools on identical speech.
The standard benchmark
The VoiceBank-DEMAND test set, published by Valentini-Botinhao and colleagues in 2016 and downloadable from the University of Edinburgh's DataShare, is 824 short sentences from two speakers, mixed with real recorded noise from cafes, streets, offices, buses and living rooms. It is the set nearly every speech enhancement paper reports on, which makes it a check on our own pipeline: the untouched noisy clips scored 1.967, matching the 1.97 in the literature.
DeepFilterNet3 scored 3.05, a gain of 1.08 over the noisy input. Its paper reports 3.17 on the same set, measured with the research code; the command-line binary at its default settings lands a little lower. RNNoise 0.2 reached 2.46, the OBS-era RNNoise 2.32, arnndn 2.15 and afftdn 2.09. DeepFilterNet3 beat RNNoise 0.2 on 784 of the 824 clips.

Averages hide the clips a tool makes worse. arnndn scored below the untouched noisy clip on 188 of 824, the OBS RNNoise on 96, afftdn on 63, RNNoise 0.2 on 57 and DeepFilterNet3 on 14. With the 12 dB cap, DeepFilterNet3 made none of them worse.
Keyboard and dog barks: noise that comes and goes
Typing and barking are the noises spectral filters handle worst, because they are not a steady background the filter can learn. afftdn left keyboard clatter exactly where it was: 1.21 before and after at 5 dB. The RNNoise models helped a little. DeepFilterNet3 took loud typing from 1.21 to 2.67, and a loud bark from 1.73 to 3.09.
The OBS RNNoise model scored 1.71 on loud barks, a hair under the untouched 1.73. RNNoise does its best work on steady noise, as the hum and hiss results below show, and a bark passes through it.
Hum and hiss
A 60 Hz ground-loop hum with its harmonics and the pink hiss of a cheap preamp are steady, which is where RNNoise works best: both versions lifted hum from 1.62 to about 2.27 and hiss from 1.23 to about 1.98. DeepFilterNet3 still led at 3.30 and 2.51. afftdn at defaults made hum slightly worse, 1.58 against 1.62.
Clean audio: the do-no-harm test
This is the result that matters if you leave a noise filter switched on all the time. We fed each tool the 100 clean sentences with no noise added. A perfect tool returns them untouched, scoring 4.64.
DeepFilterNet3 returned 4.48, and 4.52 with the cap. afftdn and arnndn stayed above 4.3. Both RNNoise versions fell to about 3.9: 46 of 100 clips from the OBS model and 58 of 100 from RNNoise 0.2 dropped below 4.0. RNNoise shapes the voice even when there is nothing to remove, so a quiet room gets a small quality penalty for nothing.
Music beds: what gets deleted
Speech enhancement models are trained to keep speech and remove everything else, and music counts as everything else. We played 80 four-second music clips, with no voice, through each tool and measured how much quieter the output was. The tracks were four Kevin MacLeod pieces from Wikimedia Commons (CC BY 3.0), mixed 10 dB under the voice for the voice-over-music test.
| Tool | Music removed (median) | Range across the four tracks |
|---|---|---|
| afftdn | 0.0 dB | 0.0 to 0.1 dB |
| RNNoise (OBS) | 2.8 dB | 1.6 to 3.0 dB |
| RNNoise 0.2 | 5.1 dB | 0.6 to 8.5 dB |
| arnndn (sh model) | 10.1 dB | 5.6 to 28.4 dB |
| DeepFilterNet3, 12 dB cap | 12.0 dB | 11.7 to 12.0 dB |
| DeepFilterNet3 | 41.8 dB | 28.6 to 52.4 dB |
41.8 dB is close to silence. Run a finished mix through DeepFilterNet3 at full strength and the voice comes out clean with the bed gone. Scored against the voice-plus-music mix a creator actually wanted, it fell to 1.30, the lowest of any tool, while the capped version held 3.54. If there is music under your voice, clean the voice track before you mix, or use the cap.

-a 12.Transcripts: does cleaning help Whisper?
Many creators denoise before generating subtitles, so we transcribed all 824 benchmark clips with faster-whisper running the small.en model, before and after each tool, and scored word error rate against the benchmark's own transcripts.
| Audio fed to Whisper | Word error rate | Clips worse than untouched | Clips better |
|---|---|---|---|
| Clean original (the ceiling) | 2.52% | ||
| Noisy, untouched | 3.75% | ||
| DeepFilterNet3, 12 dB cap | 3.70% | 32 | 24 |
| afftdn | 3.82% | 39 | 29 |
| DeepFilterNet3 | 4.18% | 46 | 30 |
| arnndn | 4.94% | 65 | 24 |
| RNNoise (OBS) | 5.18% | 77 | 21 |
| RNNoise 0.2 | 5.68% | 85 | 25 |
No tool brought the noisy audio meaningfully closer to the clean 2.52%. The capped DeepFilterNet3 tied the untouched clips (3.70% against 3.75%), full-strength DeepFilterNet3 raised errors to 4.18%, and RNNoise 0.2 reached 5.68%, about 50% more errors than doing nothing. The tool that sounds best to a person is not the one that transcribes best: Whisper already copes with background noise, and what it handled worse here were the artifacts the denoisers left behind. Transcribe the original recording and denoise the copy you publish.
Speed on one CPU core
We timed each tool on one 10-minute file (601.5 seconds of the benchmark's noisy clips joined end to end), pinned to a single CPU core, median of three runs.
| Tool | Time for 10 minutes of audio | Faster than real time | Peak memory |
|---|---|---|---|
| afftdn | 1.45 s | 415x | 24 MB |
| RNNoise 0.2, AVX2 build | 2.55 s | 236x | 3 MB |
| arnndn | 2.58 s | 234x | 24 MB |
| RNNoise 0.2, plain C build | 3.98 s | 151x | 3 MB |
| RNNoise (OBS) | 4.66 s | 129x | 2 MB |
| DeepFilterNet3 | 27.34 s | 22x | 263 MB |
Every tool runs far faster than real time on one core, so speed is not what keeps any of them off a live microphone. DeepFilterNet3 costs about 7 to 11 times the CPU of RNNoise 0.2, and it did not speed up when all six cores were available (27.78 seconds unpinned), so for a batch of files, run several at once. Its paper reports a real-time factor of 0.19 on a single-threaded notebook CPU; ours was 0.046 on one desktop core.
What OBS Actually Ships
OBS Studio's Noise Suppression filter offers RNNoise, which its documentation calls "higher quality but at the cost of greater CPU usage" than Speex. Which RNNoise? The OBS dependency scripts for Windows and macOS pin RNNoise at a snapshot labelled 2020-07-28, and on Linux OBS falls back to a copy bundled in its own source tree when no system library is found.
We compared the model weights file (rnn_data.c) in all three. They are byte-identical, and the xiph repository's history shows that file was last modified in May 2019, five years before the v0.2 release retrained the model "using only publicly available datasets". So unless your Linux distribution supplies a newer library, the RNNoise in OBS is the original model, and the column labelled "RNNoise (OBS)" in our table is a build of exactly that code.

The newer model is better on every noisy test except hum and hiss, where the two tie within 0.01, and it beat the old one on 577 of 824 benchmark clips. It is not a fix for the clean-audio penalty, which it shares. For live streaming the practical advice is the same for both: turn the filter on when there is a fan, an air conditioner or a street outside, and off when the room is quiet. On Linux, EasyEffects puts DeepFilterNet on a live microphone as its "Deep noise remover", and DeepFilterNet ships its own LADSPA plugin for PipeWire.
When Each One Wins
DeepFilterNet3 wins whenever the goal is clean speech and nothing else: podcast tracks, interview audio, voiceovers, a talking-head recording next to a keyboard. It took the top score in eight of the nine tests, and on clean input only its own capped setting did less harm. Its weaknesses are the music bed and the input format: the binary only reads 48 kHz WAV.
DeepFilterNet3 with a 12 dB cap wins when some background is part of the production: ambience in a vlog, a music bed, crowd noise at an event. It trades peak quality (2.67 against 3.05 on the benchmark) for never making a benchmark clip worse and keeping the background 12 dB below where it was.
RNNoise 0.2 is the light option for live use: 3 MB of memory and 236 times real time on one core with its AVX2 path, in a small C library under a BSD-style license. On steady hum and hiss it recovered about 40-60% of DeepFilterNet's gain; on keyboard and barks, much less.
arnndn is the one to reach for when FFmpeg is all you have in a script or on a server. It beat afftdn on eight of the nine tests, losing only on clean input, but it also scored below the untouched clip on 188 of 824 benchmark clips, the most of any tool, so listen before you ship. afftdn needs a hand-measured noise floor to do anything; see the FAQ.
How to Run Each One
All commands assume a mono voice track. DeepFilterNet's binary needs 48 kHz WAV, so convert first:
ffmpeg -i voice.m4a -ar 48000 -ac 1 voice48.wav
deep-filter -D -o cleaned/ voice48.wav # full strength
deep-filter -D -a 12 -o cleaned/ voice48.wav # keep the bed, remove 12 dBDownload deep-filter for Windows, macOS or Linux from the v0.5.6 release page. The -D flag compensates the model's delay so the cleaned file lines up with your video; without it the audio shifts. The binary loads DeepFilterNet3 by default (its README still says DeepFilterNet2; the verbose log names DeepFilterNet3_onnx).
FFmpeg's neural filter needs a model file from the rnnoise-models repository:
ffmpeg -i voice.wav -af "arnndn=m=sh.rnnn" voice_clean.wavRNNoise 0.2 builds from the release tarball with ./configure --enable-x86-rtcd && make, and its demo program works on raw 48 kHz 16-bit audio, so convert in and out with FFmpeg (-f s16le -ar 48000 -ac 1).
Pricing and Licenses
All of them cost nothing and run offline. The costs are CPU time and license terms.
| Tool | License | Commercial use | Latest release |
|---|---|---|---|
| DeepFilterNet3 (deep-filter) | MIT or Apache 2.0, your choice | Yes | v0.5.6, August 2023 |
| RNNoise | BSD-style | Yes | v0.2, April 2024 |
| FFmpeg afftdn and arnndn | LGPL or GPL, depending on the build | Yes | Part of FFmpeg |
| rnnoise-models (sh.rnnn) | Stated as not subject to copyright | Yes | 2018 models |
The cloud alternatives trade that for convenience. Our podcast tools comparison covers Adobe's free Enhance Speech, whose limits are unpublished, and Auphonic, whose free tier processes 2 hours a month. A local DeepFilterNet run has no monthly cap and never uploads a guest's voice.
Verdict
For recorded voice, run DeepFilterNet3. It was the only free tool that clearly improved every kind of noise we tried and left clean audio almost untouched. Add -a 12 the moment there is music or ambience you want to keep, because at full strength it treats your bed as noise. Keep RNNoise for live use where CPU matters, prefer the 0.2 model over the one OBS ships, and switch it off in a quiet room. Leave afftdn out of your chain unless you are prepared to measure a noise floor by hand. And for subtitles, feed the transcriber the original file, not the cleaned one; our speech-to-text comparison covers the transcription side.
How We Tested
Machine: AMD Ryzen 7 9700X, 6 virtual CPUs, no GPU. Benchmark: the 824-clip VoiceBank-DEMAND test set at 48 kHz. Creator sets: 100 clean benchmark sentences drawn with a fixed seed, mixed with keyboard typing and dog barks from the ESC-50 dataset at 5 and 15 dB, a synthetic 60 Hz hum with ten harmonics and synthetic pink hiss at 10 dB, and the music bed at 10 dB under the voice, plus the same sentences with no noise. Each tool processed each clip separately. Outputs were resampled to 16 kHz, aligned to the clean reference by cross-correlation (in our builds afftdn delayed audio by 25 ms, arnndn and RNNoise 0.2 by 10 ms) and scored with PESQ wide-band, STOI and SI-SDR. Tool versions: FFmpeg 7.0.2 static build, RNNoise 0.2 and the OBS-pinned model compiled from source with the same short driver program, deep-filter 0.5.6. Scripts and raw scores are stored in our measurements folder for this article.
Frequently Asked Questions
What is the best free AI noise removal for voice?
DeepFilterNet3, run offline with the free deep-filter binary. In our tests it scored highest on all eight noisy conditions, lifting the standard VoiceBank-DEMAND benchmark from 1.97 to 3.05 PESQ against 2.46 for RNNoise 0.2. It is MIT or Apache 2.0 licensed, so commercial use is fine.
Is the OBS noise suppression filter using the latest RNNoise?
Not going by OBS's current build scripts. They pin RNNoise at a 2020-07-28 snapshot whose model weights date from 2019, byte-identical to the copy OBS bundles for Linux builds. RNNoise 0.2, released in April 2024 with a retrained model, scored 2.46 against 2.32 for the OBS model on our benchmark.
Will AI noise removal delete my background music?
At full strength, DeepFilterNet3 removed a median 41.8 dB from music with no voice, which is close to silence. Capping it with -a 12 limits removal to 12 dB. RNNoise 0.2 removed 5.1 dB and FFmpeg's afftdn removed nothing. Denoise the voice track before mixing in music.
Why does FFmpeg afftdn barely change my audio?
Its defaults assume a noise floor of -50 dB and reduce noise by 12 dB, and on short clips it has little time to adapt. On 10 minutes of continuous benchmark audio it lifted PESQ from 2.17 to 2.27 at defaults and to 2.37 with the floor set to -30 dB and tracking on, against 3.19 for DeepFilterNet3 on the same clips. It works best when you measure the noise floor of your recording and set nf to match.
Should I denoise audio before transcribing it with Whisper?
Our test says no. Whisper small.en made 3.75% word errors on the untouched noisy clips and 3.70% after DeepFilterNet3 capped at 12 dB, a tie. Full-strength DeepFilterNet3 raised the error rate to 4.18% and RNNoise 0.2 to 5.68%. Transcribe the original recording, then denoise the version you publish.