3 ms·
How does the narration work, is it automatically generated? For a year now I have a long commute and listen to audiobooks. However I find the narration vary wil
by roywashere 3y ago
How does the narration work, is it automatically generated? For a year now I have a long commute and listen to audiobooks. However I find the narration vary wildly in quality and think oftentimes text-to-speech might actually be better
- DecoPerson 3y ago> Once we have individual tracks to work with, we begin transcription. This is the most resource intensive part of the process. We rely on the Whisper AI transcription model from OpenAI, via WhisperX. The WhisperX project also uses wave2vec2 to provide accurate word-level timestamps, which is important for sentence-level synchronization. The transcription process is fairly standard; the only interesting addition to the process that Storyteller makes is to supply an "initial prompt" to the transcription model, outlining its task as transcribing an audiobook chapter and providing a list of words from the book that don't exist in the English dictionary as hints. https://smoores.gitlab.io/storyteller/docs/how-it-works/the-algorithm/ https://smoores.gitlab.io/storyteller/docs/how-it-works/the-...
- tschumacher 3y agoYou provide an audiobook and an ebook and it syncs them.
- smoores 3y agoAs others have said, you provide the audiobook (which could technically be something you generated yourself with TTS!) and Storyteller syncs it. However, I've added an issue on GitLab to investigate building TTS directly into Storyteller, because not all books have audiobooks, and it would be cool to fill that gap!