3 ms·
On related note a very good open source TTS model was released 2 days back: https://github.com/SWivid/F5-TTS https://github.com/SWivid/F5-TTS Very good voice c
by mlboss 2y ago
On related note a very good open source TTS model was released 2 days back: https://github.com/SWivid/F5-TTS https://github.com/SWivid/F5-TTS
Very good voice cloning capability. Runs under 10G vram nvidia gpu.
- stavros 2y agoThanks! Would "under 10G" also include 8 GB, by any chance? Although I do die inside a little every time I see "install Torch for your CUDA version", because I never managed to get that working in Linux.
- mlboss 2y agoI bought a 10 Tb drive just for these kind of experiments
- linotype 2y agoTry out PopOS. They make it really easy. Though it’s named Tensorman it helps with Torch as well. https://support.system76.com/articles/tensorman/ https://support.system76.com/articles/tensorman/
- stavros 2y agoThanks, but I don't think I'm going to reinstall my entire OS to run these. I'll see if I can get Docker working, it's been more reliable with CUDA for me.
- __MatrixMan__ 2y agoI haven't tried it, but I notice that it's also in nixpkgs: https://search.nixos.org/packages?channel=24.05&show=tensorman&from=0&size=50&sort=relevance&type=packages&query=tensorman https://search.nixos.org/packages?channel=24.05&show=tensorm... That might be a less invasive way to use it, though you'd still have to install nix.
- stavros 2y agoThat's easier, thank you!
- lelag 2y agoIt actually uses less than 3 GB of VRAM. One issue is that the research code is actually loading multiple models instead of one, which is why it was initially reported you need 8 GB if VRAM. However, it cannot be used for the same use case because it’s currently very slow, so real time usage is not yet possible with the current release code, in spite of the 0.15 RTF claimed in the paper.