4 ms·
Well, no. This is a reasonable guess turned strangely confidently wrong and opinionated. Voice acting is quite literally done a sentence or at most a paragraph
by verticalscaler 2y ago
Well, no. This is a reasonable guess turned strangely confidently wrong and opinionated.
Voice acting is quite literally done a sentence or at most a paragraph at a time. Often the recording order is completely different from the script.
An actor may very well record his final scene on the first day of a project, after the whole character arc has transpired. But you know, acting. They get fed a line with stage direction and do a bunch of takes and somehow it works.
Heck you might be a full blown Italian who can't say a word in English but with the right kind of jacket it comes out a banger:
https://www.youtube.com/watch?v=-VsmF9m_Nt8 https://www.youtube.com/watch?v=-VsmF9m_Nt8
You mention Eleven labs being ahead, check out Suno. There is no LLM-scale anything involved there. The voice in this context is a musical instrument and there are lots of viable ways to tackle this problem domain.
- modeless 2y agoWe're taking about audiobooks here. An actor recording an audiobook does not read the sentences or paragraphs in a random order without context. Sure, voice acting for games or movies is done piecemeal. But the actor still gets information about the story ahead of time to inform their acting, along with their general cultural knowledge as a human. Most crucially, when acting is done in this way it is done with a human director in the loop with deep knowledge of the story and a strong vision, coaching the actor as they record each line and selecting takes afterward. When the directing is done poorly, it is pretty easy to tell. Sure, for a movie or game you could direct a TTS system line by line in the same way and select takes manually, but it would be labor intensive and not at all automatic. And to take human direction the model would need more than just the text as input. Either a special annotation language (requiring a bunch of engineering and special annotated training datasets), or preferably a general audio-to-audio model that can understand the same kind of direction a human voice actor gets.