Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
audiohermit
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
audiohermit
6y ago
Not really, this is the only thing I know of in terms of collection: https://www.isca-speech.org/archive/Interspeech_2018/pdfs/24... Usually you're basing your recipe off of those for existing datasets (
2.
▲
by
audiohermit
6y ago
Agreed. I didn't have a better comparison at hand. I'm looking at you GAN papers.
3.
▲
by
audiohermit
6y ago
Much of the work in speech synthesis has been about closing the gap in vocoders, which take a generated spectrogram and output a waveform. There's a clear gap between practical online implementations and computational behemoths like Wa
4.
▲
by
audiohermit
6y ago
I'll push back on this. The quality of the read speech should be a higher concern than having parallel data. Unless OP's wife is a teacher or actor/voice actor, if LibriSpeech transcripts are boring, it will come out in the s
5.
▲
by
audiohermit
6y ago
I work in pathological speech processing/synthesis so I'm unfortunately familiar with your father's position. It really sucks that these people didn't know that archiving their voice would've been useful. I hear sni
6.
▲
by
audiohermit
6y ago
Hey, speech ML researcher here. Make sure you have different recordings of different contexts. fifteen.ai's best TTS voices use ~90 min of utterances, some separated by emotion. If you're having her read a text, make sure it'