4 ms·
Interesting -- and unexpected. I also wrote the Speex suppressor and one of the things that specifically annoyed me about it was the robotic noise and the pseud
by jmvalin 9y ago
Interesting -- and unexpected. I also wrote the Speex suppressor and one of the things that specifically annoyed me about it was the robotic noise and the pseudo-reverberation it adds to the speech, but it seems like some people (like you) like that. Trying to understand exactly what you don't like about RNNoise... is it how the remaining background noise sounds or how sharply it turns on/off?
I did a quick hack to RNNoise to smooth out the attenuation and prevent it from cancelling more than 30 dB. I'd be curious if it improves or makes things worse for you (compared to the samples in the demo):
https://jmvalin.ca/misc_stuff/rnn_hack1/ https://jmvalin.ca/misc_stuff/rnn_hack1/
- mrob 9y agoI also much prefer the Speex version. The robotic noise is consistent and easy to ignore. The NN version has a "choppy" feel to it that catches my attention. It reminds me of H264 vs VP9 video codecs at the same bitrate. VP9 is supposed to be better, but H264 artifacts are more artificial looking. VP9 artifacts look more natural, almost like mold/decay on film, which I find more distracting. I prefer your hacked RNNoise version to the original, but I still prefer the Speex version. Robotic is predictable, and predictable is good. I don't want denoising artifacts to feel like there's some intelligent agent behind them, just as I don't want any software tool to feel intelligent. The smarter the tool the more jarring it is when it misreads my intentions. It might help average performance but it harms worst-case performance, and worst-case performance is subjectively more important because humans pay attention to outliers.
- mtrimpe 9y agoSubjectively it seems to me that the RNNoise sample doesn't trigger my brain to attempt to fill in the gaps. With the Speex/raw ones I have all the data so if I listen to it again over and over I can get more out of it eventually. With the RNNoise one I obviously don't even have enough extra data to even try doing that so all I can do is blame the algorithm. Perhaps what you really want is an algorithm that lets through a bit more of the 'possible noise' for the human brain to have another go at.
- jmvalin 9y agoWhat you're describing is more or less why noise suppression algorithms in general cannot really improve intelligibility of the speech. Unless they're given extra cues (like with a microphone array), there's nothing they can do in real-time that will beat what the brain is capable of with "delayed decision" (sometimes you'll only understand a word 1-2 seconds after it's spoken). So the goal of noise suppression is really just making the speech less annoying when the SNR is high enough not to affect intelligibility. That being said, I still have control over the tradeoffs the algorithm makes by changing the loss function, i.e. how different kinds of mistakes are penalized.
- mtrimpe 9y agoPerhaps being more lenient in noisier situations could be an interesting tradeoff then. At lower noise levels it's already pretty good...
- Jasper_ 9y agoI think the biggest issue for me was the sharpness of the cutoff -- it actually sounded like a simple noise gate to me. The hack here helps a lot in smoothing out the sharp attack.
- aidenn0 9y agoI found terminal sounds, fricatives and sibilants to be at a mininal distracting with RNN for "car" and "street", and at worst unintelligable. In particular, the terminal sound in "Christmas" was completely lost in the noise for me at 10dB with RNN, but was perfectly fine in Speex. For "car" RNN sounded as good or better than Speex at all noise levels. That rnn_hack is significantly better for me. 5dB on that sounds strictly better than 10dB on the original to my ear for "babble" and "street". I also noticed that the for the parts that sound the worst to me at 10-15dB in the original RNN, the signal is completely missing in the 0dB RNN version, so perhaps the signal is in the same band as the noise at that part? Either way it's a tough tradeoff because I suspect that low bitrate encodings will love the nearly empty signal in the bands that are generated by the original, but the seemingly rectangular cutoff/introduction of the noise was much more jarring to me than the reverberation added by Speex (though I didn't like that in Speex, it didn't seem to add to my effort to understand the way that.
- grandalf 9y agoFWIW I much prefer the RNNoise version to the Speex version. With the Speex version the nosie is much more consistently present/noticeable/distracting.
- belthesar 9y agoIt reminds me of the introduction of line noise into VoIP systems to replicate the natural electrical noise present in POTS systems. Without something, the hard attenuation brings more attention to the noise that still exists.