3 ms·
While we’ve had rnnoise integration for a while it was for “noisy environment” notifications, this is the first time we use it to actually filter audio. Also a
by saghul 4y ago
While we’ve had rnnoise integration for a while it was for “noisy environment” notifications, this is the first time we use it to actually filter audio.
Also audio worklets weren’t a thing when we first introduced it.
I’m not aware of any other open source (and better) models, but if any come up, we’ll certainly check them out!
- pen2l 4y agoIf you have any involvement with Jigasi or might be in the know -- are there plans to use whisper, for instance, instead of Google's API for transcription? If I recall correctly jigasi is using google's API, local transcription aligns well with the rest of Jitsi's missions.
- saghul 4y agoWe do have VOSK support already. I haven’t heard of whisper, but it does sound like a good GSoC project for next year!
- pen2l 4y agoIf I have time I'll try to help you guys out. I'm a big fan of what you're doing. :)
- nikvaes 4y agoThe problem for Jigasi's speech-to-text feature with Whisper - or any recent SOTA speech-to-text neural networks, is that they are transformer-based. One of the key features of transformers is that they are very good at processing a sequence with the attention mechanism. But attention inherently needs to see the whole input sequence. So it's difficult to adapt these architectures to perform well in real-time scenarios like captioning meetings.
- pen2l 4y agoYes! But a part of the Jitsi ecosystem enables recordings and whisper is a good candidate to use for these recorded sessions. On that topic — they record sessions in an interesting way, basically an instance of chrome is started and captured... I think with OBS. That always made me raise an eye but I also can’t think of up a better way. edit: It's actually jibri which has to do with recording. Gosh I wish the names were a liiiittle more intuitive. :)
- eis 4y agoThanks for the clarification. We also experimented with audio worklets + rrnoise about 1.5 years or so ago but had very mixed results. The potential upside with processing in another thread is clear but some browser and OS combinations just didn't work well and resulted in micro stutters in the audio. I remember Chromium on Linux for example being finicky. Some browsers worked better with smaller buffers, some needed bigger ones. We spent too much time debugging and tuning for different systems and the audio quality improvement was not deemed good enough so we shelved the effort. I guess audio worklets improved since then and probably is more useable by now. Do you guys have some kind of performance monitoring for the noise cancellation or audio in general? At the time I also spent a few days looking for something better but didn't really find anything. Unfortunately RRNoise is the best we have :( The only other noise cancellation software that actually impressed me was the one from Nvidia but that's not something that one could integrate via WASM and of course wouldn't work on most devices anyways. Oh what a day it will be where we have energy efficient hardware encoders for AV1 in every device plus some really good noise cancellation. Oh and then we just need internet connections without packetloss :P