2 ms·
According to the documentation[1], it's a concatenative synthesizer using decision trees for prosody modeling and PSOLA for output. [1]: http://www.ibm.com/sma
by cypher543 12y ago
According to the documentation[1], it's a concatenative synthesizer using decision trees for prosody modeling and PSOLA for output.
[1]: http://www.ibm.com/smarterplanet/us/en/ibmwatson/developercloud/doc/text-to-speech/#research http://www.ibm.com/smarterplanet/us/en/ibmwatson/developercl...
- kastnerkyle 12y agoThanks! I am working in this area and have some ideas for deep learning type methods which move away from concatenative synthesis. It will be nice to compare to what they are using.
- woodson 12y agoThis paper (from ICASSP2013) may be of interest to you: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/40837.pdf https://static.googleusercontent.com/media/research.google.c...
- picheny 12y agoWe did some work on applying NNs to prosody prediction; see Fernandez, Raul, et al. "Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networks." Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH). 2014.