3 ms·
If this lives up to the demo it's a huge development for anyone looking to do realistic tts without paying to use an API
by nopelynopington 1y ago
If this lives up to the demo it's a huge development for anyone looking to do realistic tts without paying to use an API
- kristopolous 1y agothere's quite a number of pretty low overhead models around that do that in realtime these days.
- foofoo12 1y agono
- MarsIronPI 1y agoBut how many of them support voice cloning? (Genuine question; I haven't seen any other than this one.)
- nickthegreek 1y agomicrosoft’s vibe voice.
- MarsIronPI 1y agoVibeVoice (according to the repo description) is currently unavailable due to "misuse". But my impression was that it required a significant (>8GB) amount of VRAM? Or that it wasn't suitable for on-device for devices with low specs.
- nickthegreek 1y agoits unavailable from their repo, but was released with an open license and mirrors exist. I'm not sure what the VRAM req are.
- MarsIronPI 1y agoAccording to this issue[0] the 1.5B model needs 6GB of VRAM. Meanwhile it looks like NeuTTS is designed to be able to run on CPU, which is nice for older/lower-spec hardware. 0: https://github.com/microsoft/VibeVoice/issues/26#issuecomment-3227465614 https://github.com/microsoft/VibeVoice/issues/26#issuecommen...