3 ms·
Subjectively it seems to me that the RNNoise sample doesn't trigger my brain to attempt to fill in the gaps. With the Speex/raw ones I have all the data so if
by mtrimpe 9y ago
Subjectively it seems to me that the RNNoise sample doesn't trigger my brain to attempt to fill in the gaps.
With the Speex/raw ones I have all the data so if I listen to it again over and over I can get more out of it eventually.
With the RNNoise one I obviously don't even have enough extra data to even try doing that so all I can do is blame the algorithm.
Perhaps what you really want is an algorithm that lets through a bit more of the 'possible noise' for the human brain to have another go at.
- jmvalin 9y agoWhat you're describing is more or less why noise suppression algorithms in general cannot really improve intelligibility of the speech. Unless they're given extra cues (like with a microphone array), there's nothing they can do in real-time that will beat what the brain is capable of with "delayed decision" (sometimes you'll only understand a word 1-2 seconds after it's spoken). So the goal of noise suppression is really just making the speech less annoying when the SNR is high enough not to affect intelligibility. That being said, I still have control over the tradeoffs the algorithm makes by changing the loss function, i.e. how different kinds of mistakes are penalized.
- mtrimpe 9y agoPerhaps being more lenient in noisier situations could be an interesting tradeoff then. At lower noise levels it's already pretty good...