4 ms·
LibriSpeech, Tatoeba, Common Voice and scraped YouTube videos.
by iceychris 6y ago
LibriSpeech, Tatoeba, Common Voice and scraped YouTube videos.
- blackcat201 6y agoDo you get good results when adding scraped youtube audio? My model performance on LibriSpeech dev drops a bit when adding youtube audio to the training dataset ( my guess is likely due to poor alignment from auto generated captions ).
- iceychris 6y agoI haven't trained on LibriSpeech exclusively, but yes, the perf on LibriSpeech dev is quite bad, around ~60.0 WER. If the poor alignment of yt captions is the issue, maybe concatenating multiple samples helps a bit.
- lunixbochs 6y agoYou should consider realignment; maybe start with something like DSAlign or my wav2train project.
- dcsan 6y agowould it be possible to train on any of the more recent Text to speech engines out there? some of them are very realistic. this would give you absolutely perfect sync down to the word, I assume... I don't know about the cost if you paid ratecard though, perhaps you can do some partnership with them since yours is a symetrical product