4 ms·
>So how long will it be until we will be able to download something of this quality onto a future-gen Raspberry Pi which can do some AI processing, where we mak
by JonathanFly 3y ago
>So how long will it be until we will be able to download something of this quality onto a future-gen Raspberry Pi which can do some AI processing, where we make an HTTP call and it starts speaking through the audio out in a perfect voice without relying on the cloud?
5 years? It's probably possible roughly whenever the larger Whisper models can run on it. Probably the next Raspberry Pi, running quantized or optimized versions of some audio model.
It may be almost possible right now if you tried really realy hard, and you used a small model fine-tuned on a single voice, instead of something larger and more general purpose that can do any voice. I think whisper-tiny works on a Pi on real time, right? And that's not leveraging the GPU on the Pi. (https://github.com/ggerganov/whisper.cpp/discussions/166 https://github.com/ggerganov/whisper.cpp/discussions/166)
Edit: looks like medium is 30x slower on the Pi than tiny model, so I may have been overly optimistic. I didn't realize Whisper tiny was that much faster than medium.
This method works pretty well with Tortoise, letting you use the super fast Tortoise quality settings but get quality similar to the larger models. Fine-tuning the whole thing on just one voice removes a lot of the cool capabilities of course. With Tortoise, that would still be way too slow for a Pi but potentially that same strategy could work with faster models like SoundStorm.
In terms of quality there's still a lot of room to go with long term coherence, like long audio segments. When a real person reads an audiobook the words at the top the page have a pretty big impact on how many words at the bottom the page are read. And there can be some impact at any distance, page 10 to page 300. When you try audiobooks on super high end TTS models and listen carefully you really notice the mismatch. It's like the reader recorded the paragraphs out of order, or a video game voice lines where you can tell the actors recorded all the lines separately, and were not reacting to each other's performance.
You can bump the context windows, a minute, two minutes. That's gonna get you closer and probably good enough for some books. In the short term a human could simply adjust all the all the audio samples and manually tweak things to sound correct. So this will enable fan-created audiobooks where they take the time to get it right. But for fully automated books the mismatch drives me nuts. The performance is just soooo close for certain segments that when you get a tonal mismatch it hurts.