4 ms·
> An ideal dataset would consist of thousands of hours of speech where source accent utterance is mapped to each target accent utterance and aligned with it acc
by bittlingmayer 3y ago
> An ideal dataset would consist of thousands of hours of speech where source accent utterance is mapped to each target accent utterance and aligned with it accurately.
To put it in terms of text translation, roughly how many sentences or words is this?
- bittlingmayer 3y agoPardon me, how many "tokens" ;-)
- davitb 3y agoOn average, an hour of speech contains about 9,000 to 15,000 words. This range accounts for different speaking speeds, which typically vary from 150 to 250 words per minute. So this translates to tens of millions of words.