3 ms·
I would say speech synthesis is governed by very clear mechanics. https://en.wikipedia.org/wiki/International_Phonetic_Alphabet https://en.wikipedia.org/wiki/In
by dbelchamber 8y ago
I would say speech synthesis is governed by very clear mechanics.
https://en.wikipedia.org/wiki/International_Phonetic_Alphabet https://en.wikipedia.org/wiki/International_Phonetic_Alphabe...
As for style transfer, that is a very specific skill of making the patterns of one style map to the patterns of another. I am not particularly well versed with art, but that process seems well defined to me.
Perhaps your issue is with my more generalized definition of "clear" and "well-defined". I meant to use these terms to distinguish between autonomous driving and being a successful human. I really don't think there is anywhere close to a consensus on the latter. To the extent that there is, then yes, AI should be able to do it.
- Veedrac 8y agoThe IPA is nowhere close to sufficient for realistic speech synthesis, and style transfer is not just copy and paste. By the same token writing poetry is just "putting words into grammatical constructions that have certain patterns" or mathematics research is just "a form of proof search". Of course we don't have human-level AI right now, but if that's the only thing you're claiming it's pretty vacuous.
- dnautics 8y agoI would say speech synthesis is governed by clear mechanics - and it's not the IPA, it's that the output comes out as a waveform, which has a structure that informs the algorithm. Note that we have great raster-based deep visual effects, but vector is... not there yet (not saying it won't be) - vector is less structured than raster, so the choice of algorithm is less obvious. As for well-defined criteria, I don't think that's really quite the right standard, I think the correct standard is that there is a way of metrizing success on a well-ordered set (like the [0,1] interval), even if it's noisy.