6 ms·
My 26 second training input perhaps wasn't enough. The result sounded like someone else. Is the result some kind of merger of my voice and a native speaker's?
by sxv 5y ago
My 26 second training input perhaps wasn't enough. The result sounded like someone else. Is the result some kind of merger of my voice and a native speaker's?
- reubenmorais 5y agoSimilarity depends on many factors: recording quality, which language you're synthesizing in (models trained on more speakers do better), and diversity of prosody in your recording. Try recording for a bit longer and "acting out" a bit in your tone, that tends to give me interesting results :)