3 ms·
How did u manage the sync between the audio and words? ... In my little demo, I've used IBM Watson's Speech to text API to analyze the audio. I used those resu
by avantion 9y ago
How did u manage the sync between the audio and words? ... In my little demo, I've used IBM Watson's Speech to text API to analyze the audio. I used those results - it had timestamps, to the recognized words - and used some algorithms to match it with the text. Naturally, it did not catch like 20% of the words but was able to massage it enough that it looked smooth.
- vitovito 9y agoI had a verbatim transcript already, and I used CMU Sphinx to do the alignment between that and the recording. With raw audio only, I'd probably use a transcription service that provides timings, or software like Trint to do machine learning first, and then clean up by hand, like you did. The software I wrote to handle the playback is here: https://github.com/vitorio/hyperaudio-lite https://github.com/vitorio/hyperaudio-lite