4 ms·
I've been following jpc [0] on the LAION discord since he started building this last year, and it's a very impressive project. The key here is that the Whisper
by nmfisher 3y ago
I've been following jpc [0] on the LAION discord since he started building this last year, and it's a very impressive project.
The key here is that the Whisper multilingual ASR model has been trained on a huge amount of data, so its encoder output is a very good representation of the semantic content of speech. This can be used as an open-source, drop-in replacement for the semantic encoder in model architectures like SPEAR-TTS/VALL-E/etc (whose semantic encoders are not publicly available). This is then used to predict acoustic tokens (the output from the quantized/low-bandwidth Encodec audio codec) which is then upsampled/denoised/enhanced with the Vocos vocoder.
I know someone is working on Hindi but it would be great to see this extended to other languages for a properly open-source [1], multilingual TTS platform. I think the main bottleneck at the moment is finding people who can procure/clean compliant datasets.
[0] https://github.com/jpc https://github.com/jpc
[1] jpc/Collabora went to great efforts to ensure that they are only using properly licensed data to train this. I doubt Whisper itself was that compliant, so it's a bit muddy.
- 3abiton 3y agoThey don't mention the ability to add custom voices to the speech output, I wonder if that's a feature thatbwould be supported
- atwrk 3y agoThey do mention voice cloning in the README ("We’ve also added an example of voice cloning based on a reference audio file."), do you have something different in mind?
- addandsubtract 3y agoIt's only found in the Colab and not in the Readme, though. The examples in the Colab are also better than the ones found in the Readme. Maybe the Readme still needs to be updated?
- jpcl 3y agoHi, thanks a lot for the tip, I'll update the README samples ASAP. :) I was busy working on inference performance in the last few weeks and totally did not expect to land on Hackernews today. Only noticed it an hour ago because my GitHub stars jumped quite a bit
- jpcl 3y agoYeah, Whisper is not clear-cut but since it is not a generative model I think their data usage is a lot more likely to be considered fair-use. And the part of that which we use for WhisperSpeech is just the phonetic representation so our model is not able to recreate any of the Whisper training data in any way.
- leereeves 3y agoThe readme says "We are working only with properly licensed speech recordings and all the code is Open Source so the model will be always safe to use for commercial applications." Is that less certain than the quote implies?
- jpcl 3y agoWe are working hard to uphold all the licensing rules but nobody can absolve you from all legal risks. There may be a court ruling/new law that any training needs a special permission from the original author and then even a CC-BY license won't cover this.
- doctorpangloss 3y agoLaypeople value the aesthetics of statements like these. It's very Discord energy. Everyone using learned weights from other models, especially ones released by OpenAI, Stability and Google, such as text and audio encoders, is tainted by training materials that were not expressly licensed for the purpose of AI model training or unlimitedly licensed for any use.
- jpcl 3y agoThat's true but you make it sound like it's totally obvious where the line of fair use should be drawn for AI training. Until courts or lawmakers make it clearer I personally believe non-generative models (Whisper, ResNet, DINOv2) should be legally trainable on publicly released data. Generative models (image or video generation, TTS, LLMs?) should be held to a much higher scrutiny since their outputs can potentially compete the creators who put a lot of creativity into their art. That not true for an ImageNet-trained classification model or Whisper ASR.