2 ms·
New multimodal models take raw speech input and provide raw speech output, no tts in the middle.
by computerex 2y ago
New multimodal models take raw speech input and provide raw speech output, no tts in the middle.
- benob 2y agoA relatively detailed description of such systems: https://arxiv.org/abs/2410.00037 https://arxiv.org/abs/2410.00037
- Closi 2y agoSeems like the future - so much meaning and context is lost otherwise.
- intalentive 2y agoVery cool. Logical next step. Would be interested to know what the dataset looks like.
- moffkalast 2y agoYoutube. Youtube is the dataset.