4 ms·
If you have any involvement with Jigasi or might be in the know -- are there plans to use whisper, for instance, instead of Google's API for transcription? If I
by pen2l 4y ago
If you have any involvement with Jigasi or might be in the know -- are there plans to use whisper, for instance, instead of Google's API for transcription? If I recall correctly jigasi is using google's API, local transcription aligns well with the rest of Jitsi's missions.
- saghul 4y agoWe do have VOSK support already. I haven’t heard of whisper, but it does sound like a good GSoC project for next year!
- pen2l 4y agoIf I have time I'll try to help you guys out. I'm a big fan of what you're doing. :)
- nikvaes 4y agoThe problem for Jigasi's speech-to-text feature with Whisper - or any recent SOTA speech-to-text neural networks, is that they are transformer-based. One of the key features of transformers is that they are very good at processing a sequence with the attention mechanism. But attention inherently needs to see the whole input sequence. So it's difficult to adapt these architectures to perform well in real-time scenarios like captioning meetings.
- pen2l 4y agoYes! But a part of the Jitsi ecosystem enables recordings and whisper is a good candidate to use for these recorded sessions. On that topic — they record sessions in an interesting way, basically an instance of chrome is started and captured... I think with OBS. That always made me raise an eye but I also can’t think of up a better way. edit: It's actually jibri which has to do with recording. Gosh I wish the names were a liiiittle more intuitive. :)