6 ms·
I'm having a hard time coming up with a non-nefarious use case for this.
by jascii 4y ago
I'm having a hard time coming up with a non-nefarious use case for this.
- jeroenhd 4y agoVoice generator tech has created some decent surreal memes (like audio recordings of Biden, Obama, and Trump playing video games together). Outside of memes or maybe the occasional well-intentioned prank, I really can't think of anything either.
- Rubinsalamander 4y agoMassively reducing costs for Voice Over in Video Games. This should make it even feasible to create mods with audio which would be great :)
- inerte 4y agoI think “talking” with dead relatives or friends will become real pretty soon. If people can find comfort hearing their mom say words of encouragement in a tough situation, I think a lot of people would do it. Kinda hard because for some others that would mean never getting closure. Weird stuff is certainly about to happen…
- woodrowbarlow 4y agothere has been some media coverage on this already (e.g. [1]). an emerging concern among mental healthcare professionals is that a sufficiently-convincing simulation could interfere with the progression of the stages of grief, prolonging the 'denial' stage and potentially heightening the intensity of the stages that follow. [1] https://www.wired.com/story/a-sons-race-to-give-his-dying-father-artificial-immortality/ https://www.wired.com/story/a-sons-race-to-give-his-dying-fa...
- starkparker 4y agoThe last thing on earth I'd want is for any aspect of my dead relatives to be reanimated through technology. No. That's absolutely fucking horrific to consider. I don't need a hallucinating AI pretending to be my dead wife. That's literally shambolic. There is vastly more potential for that to be abused by others than used in any emotionally or socially constructive way.
- Rubinsalamander 4y agoI would also find that very creepy and it would probably keep you from moving on. I think there is a big difference between remembering what happened by looking at a photo or hearing an audio recording and having newly generated "content" from a deceased loved one.
- jeroenhd 4y agoI would consider studios taking voice actors' voices and using them to generate new content beyond their contract to be abuse. I'm sure big corporations are rubbing their hands in anticipation, but I'm sure killing the VA industry will make the world just a tiny bit worse for everyone else. Mods are more difficult to attach a moral judgement to. I don't think I'd really consider them malicious, as long as they're not sold, but there's a very thin line between a high quality mod and stealing someone's voice.
- Izkata 4y agoStar Trek: Prodigy has already used audio from previous movies and TV to bring back to life several actors from previous series. It's not exactly the same as this, but their dialogue was taken out of context to create new scenes and story.
- jeroenhd 4y agoI know, and I almost wished they did use AI for that segment because it was pretty jarring (especially the TOS recordings). There's still a huge difference between "reusing the work the studio paid for" and "recreating your voice forever after doing a single project".
- buu700 4y agoOn the other hand, why shouldn't voice actors benefit from this tech? I can easily imagine a future where AI-generated impersonations are deemed by courts or new legislation to be protected by personality rights. In that world, voice actors could expand their business by offering deeply discounted rates for AI-generated work. Alternatively, if/when tech like Play.ht is consistently good enough, maybe it just becomes a standard practice for all voice acting work to include a combination of human- and AI-generated content, like a programmer using Copilot or a writer using GPT.
- gamblor956 4y agoI'm sure programmers would love to expand their business opportunities by offering deeply discounted rates for creating AI-generated code. No? Then why do you assume that someone else would want to do the same in their profession? As AI-generated content is not protectable under IP law, it's a non-starter for games, film, TV, or music for anything except background filler.
- rockemsockem 4y agoAnything written can be listened to with this tech. Any news article, any short story, a draft of a piece of writing you're working on. There is too much text for human beings to read it all.
- scrollaway 4y ago> There is too much text for human beings to read it all. so your logic is that all that text should be audio and people will consume more? Because I got news for you, reading is faster than listening.
- throwaway675309 4y agoOh yeah, how does reading work out for you while you're driving a car? smh...
- rockemsockem 4y agoWhen I said there's too much text for human beings to read it all I meant that it isn't feasible to pay people to read all text that someone might want to listen to into audio. Like a random blog written by someone in their spare time probably isn't going to hire a voice actor. I think the case for having all text be listenable is pretty clear. We're all really busy and often our hands are busy but we're not doing something that mentally stimulating. This is an ideal time to listen to an audiobook, a blog, the news, or whatever else you'd like.
- vincnetas 4y agoAnd all AI bots are here to generate even more text. :( We will need to rethink and reevaluate lots of things that we are used to.
- bovermyer 4y agoI'd get a kick out of having my own blog posts read to me in James Earl Jones's voice. Or, heck, my own voice. Though it'd be surreal to hear not-me-but-me saying things I've never said.
- woodrowbarlow 4y agoeven this is ethically questionable. james earl jones's voice is his livelihood.
- bovermyer 4y agoWhile that is true, I'm not suggesting a pattern of behavior - just that it would be fun to hear.
- atentaten 4y agoGenerating audio for an audio book: If an author could speak for 20 minutes and then generate audio for an entire book from the book's text and the model, I think that would be very useful.
- sva_ 4y ago20 seconds*
- atentaten 4y agoThe OP mentioned that for so called, "High-fidelity voice cloning", it would take 20 minutes of training. I think a book author would want the best quality possible to reproduce their voice.
- JohnFen 4y agoWhy reproduce their voice? There's no value-add there.
- sva_ 4y agoMany people prefer an audiobook version of a book to be read by the original author, which isn't always the case. If an author could make that version happen by using 20 minutes of their time + text2speech of the whole book, that would be an immensely positive value proposition on the side of this company. But I'm not sure. Part of why I'd prefer the original author to read a book is that they vocally emphasize certain parts of the book, and I don't think these models could do that at this point.
- JohnFen 4y ago> Many people prefer an audiobook version of a book to be read by the original author Right, but having AI read the book in the author's voice is definitely not the author reading the work. As you mention, the reason that people like to hear the author read it is because it's the author reading it, theoretically emphasizing and acting things out according to what was intended. It's not just to hear the author's voice. So I don't see what the value-add is.
- zanderwohl 4y agoI am toying about with building a virtual puppet software in the style of watchmeforever. I have a number of voices I do for the stage and DnD that I would be willing to train a few models on so I could give my puppets unique voices.
- mahmoudfelfel 4y agoWe have been seeing some of these genuine use cases: youtube creators, audiobooks, elearning videos, podcasts, commercials, dubbing, and gaming.
- erichocean 4y agoI'm using this kind of technology for temporary voice tracks in animated shorts. I'd really like something like Img2Img for voices so I can translate a performance to an arbitrary (synthetic) voice.
- nullsense 4y agoTortoise TTS can do this. You just pass it your example as a conditioning latent.
- erichocean 4y agoThanks!