3 ms·
For this zero-shot approach, there's actually no "fine-tuning", and from what we can tell ~30 seconds is optimal (the text isn't actually even used... you can s
by _josh_meyer_ 4y ago
For this zero-shot approach, there's actually no "fine-tuning", and from what we can tell ~30 seconds is optimal (the text isn't actually even used... you can say anything!)
For longer outputs and legit fine-tuning, that could become an offering :D
- windsignaling 4y agoIs it possible to do the same thing (zero-shot on an arbitrary input voice) using one of the models on your Github?