10 ms·
Voicebox: Generative AI model for speech that generalizes across tasks
- yewenjie 3y agoPretty interesting. IIUC they haven't released the model weights, but I think most TTS papers in the upcoming months will release the weights since OpenAI Whisper.
- canadianwriter 3y agoThat little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
- Kiro 3y agoWhat stinger at the end?
- az226 3y agoNot releasing code or weights under the false pretense of misuse.
- JimDabell 3y agoNo, they are referring to the end of the video, where they have a “surprise twist” that the video narration was autogenerated and not a real person.
- az226 3y agoI mean that wasn’t even slightly non-obvious.
- flangola7 3y agoFalse? The misuse opportunities are obvious.
- az226 3y agoReproducing the code is a matter of time, and short at that.
- emehex 3y agoI agree with you that v1 isn't a suitable replacement... right now. But what v2? v3?
- graypegg 3y agoI don’t think GP is saying it won’t improve. The thread is about the current state of it that meta wrote this article on.
- deleted 3y ago[deleted]
- twelfthnight 3y agoSPOILER ALERT :p Its more surprising nowadays if an article about AI doesn't have a "twist" that what you read/heard/saw was AI.
- pixl97 3y agoFor me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.
- chrislikescode 3y ago[flagged]
- bloqs 3y ago[flagged]
- ignoramous 3y agoCryptographic proof of personhood is going to be a thing, is it not? Outside of BigTech, Signal is as poised as WorldCoin to be just that.
- waboremo 3y agoYes and it's going to be done through digital IDs. Unless something dramatic happens, we're poised to turn to digital IDs linked to your real ID and in turn validating access to apps/communication.
- pixl97 3y agoAnd authoritarians everywhere will rejoice (and they will give out the means to duplicate these IDs to a select few in case they need to generate evidence that 'you' have offended the state).
- kybernetikos 3y agoThe only sensible approach to this problem (assuming it is a real problem) is trusted individuals certifying others as human. There are alternatives to using the state for this, but they are difficult and fraught with UX issues. Perhaps a decentralised web of trust or some sort of blockchain based registrar of trust that can trace trust routes between mutually distrusting individuals. Unless such a system is in place, international and strong before states start playing in this space, there isn't much chance of beating a state's approach to the problem. Just look at https certificates. The current system involves browsers shipping configured to trust a whole bunch of entities I don't really trust, and there has been relatively little interest in trying to build a working decentralised approach to site security.
- ethbr0 3y agoFor video narration elocution, I'd say it was most of the way there. When. Narrating. Videos. One. Tends. To. Speak. Differently. Or, the more important case -- if I'm listening to audio-version-of-X, is it sufficiently human-like that I can forget that it's synthesized voice? To me, yes. Easy to tell if you're specifically listening for it, but to use an analogy one doesn't typically read novels and parse closely for grammar, does one? Your attention is elsewhere, on the content and plot.
- TacticalCoder 3y ago> It has a very obvious "autotune" To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work...". Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech. The first seven seconds are particularly bad. Don't get me wrong: I've got a lovely french accent when I speak english. This has either been trained on too many audiobooks spoken by non-natives or they've used their own tech, where the "reference audio" given as input was from a non-native. In any case something is seriously off. At 1:59, the "Hi guys, thanks you for tuning in! Today we are going to show you..."... That is obviously an Hindi speaker speaking (it's an example of fixing a real voice by removing background sounds). I think that the main voice of the video was done by the same person who did the example at 1:59. And I think that they used their example of using a "reference audio". And that person ain't a native english speaker. To compare: when the reference audio uses a proper english accent (the example with the "diverse ecosystem" at 0:52), then the output from the text-to-speech sounds native. I think they just fucked the demo video and it may already be ready for prime time.
- kybernetikos 3y agoMaybe they deliberately chose an accent that wasn't native English to demonstrate the style transfer capability. I think the ability of the system to output accented voices is a strength not a weakness, so long as it can do other accents too.
- bee_rider 3y agoThe accent was obvious enough that I wonder if they might have not been trying to hide it at all? Maybe they just happened to pick somebody from the team with a very mild accent.
- alach11 3y agoI'm surprised you had such a negative reaction to the Hindi accent! To me, it was no more difficult to understand than my colleagues who speak English as a second language. To me, this is a style choice for the demo. Not evidence that they "fucked" it up. Accents are common - everyone has one! It's nice to see the model can support your personal voice even if it's not completely neutral English.
- aketchum 3y agoI think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates
- deleted 3y ago[deleted]
- canadianwriter 3y agolike... a pin?
- madsbuch 3y agoIn the US some phone companies have been using voice recognition to authenticate when their customers call. This will definitely have to see its end.
- zimpenfish 3y agoHMRC in the UK also use it. Has never worked once for me and I don't even have any kind of accent.
- bool3max 3y agoHow "live" can translation ever really be? Properly translating anything from one language to another requires context.
- bugglebeetle 3y agoWhile I’m not necessarily in favor of this, a multimodal AI that has access to your location, vision inputs, etc could obtain much of this context. People already explore foreign countries by hobbling together these services.
- 3y ago
- ChatGTP 3y agoIs Meta just operating in hope and AI moonshots now ? Everyone one of their products is just garbage to me and becoming less relevant by the day. When do they actually starting building something useful again ? Honestly Apple seems to be using “AI” much more successfully and actually seamlessly integrating it into their existing products to improve them. My theory is Mark is hoping the meta verse will pop out if Yan’s bottom at some stage. Maybe he is right? I just can’t for the life of my understand why the current products are just so so neglected?
- ChuckNorris89 3y ago[flagged]
- ChatGTP 3y agoDo you use an iPhone ? They do pretty amazing things with images now. Even the search for a photo by text is quite amazing. I used it the other day for work for the first time and I found what I needed in my tens of thousands of photos. It almost truly felt like an extension of my memory, it was actually a pretty cool feeling. So sorry I don’t buy the Siri attack as being proof of anything. I have found Siri has improved a lot. Even the speed at which Siri works is much faster. The other day I was driving in a really loud old car and used my Apple watch to change the music, I said to myself “there’s no way Siri will get it”, and it did. I also trust Apple with my data, at least for now. I don’t even slightly trust Meta at all. I found Zuckerbergs interview on Lex Friedman more scary than before as he said all the same freaky stuff disguised by a more “cool” and polished facade. The same dystopian ideas are still there.
- ChuckNorris89 3y agoSent from my iPhone
- deleted 3y ago[deleted]
- twobitshifter 3y ago
- cnlwsu 3y agoI am mostly excited for cheaper audiobooks with consistent voices for different characters.
- mjamesaustin 3y agoOh that is an amazing use case!
- benabbottnz 3y agoWhy would they make it cheaper when they can make even more profit by not having to pay a voice actor?
- trts 3y agoIf they make it better for the reader, they can potentially raise the price. If they can make it cheaper to produce, they can potentially increase their profit without raising the price. Usually on balance this falls somewhere in between -- more value for less money for the consumer, and more profit on each marginal unit of production for the producer, which is how technology progresses across most consumer goods.
- madsbuch 3y agoBecause it would be extremely easy to produce it cheaper or record it independently. This opens up for non-signed authors to release audio books.
- waboremo 3y agoYes, but they're not going to pass on costs from not using a voice actor. They're just going to charge what they normally would have, and not worry about having to give so and so a cut.
- ericd 3y agoSure, and then we'll get more audiobook options, as it becomes economically viable to make more niche stuff into audiobooks.
- rvz 3y agoAnd the source code for this one has.... ...not been open sourced and cannot be found. Sorry AI bros. Better read the paper this time.
- gwern 3y agoIt won't be a big barrier. Voice stuff is not as computationally intense as video or LLMs, so it's still an area where small teams or hobbyists can make a dent.
- IshKebab 3y ago> There are many exciting use cases for generative speech models, but because of the risks of misuse, we are not making the Voicebox model or code publicly available at this time. While we believe it is important to be open with the AI community and to share our research to advance the state of the art in AI, it’s also necessary to strike the right balance between openness with responsibility. Yeah... I mean there definitely are ways to misuse this (especially the style transfer!) I don't think you're going to do anything except delay the inevitable Facebook.
- sgift 3y agoThey don't really care about misuse. They just don't want to say openly that they like to keep their shiny new tech and make money with it. No idea why, most people wouldn't bat an eye if you stated from the get go: We built it, we'll use it.
- IshKebab 3y agoI agree. Though in fairness they have opened other models e.g. for speech recognition.
- swader999 3y agoSo where can we try this out?
- theptip 3y ago> See voicebox.metademolab.com for a demo of the model They are not releasing the model (yet?) but demo samples are available across many tasks
- moneywoes 3y agoIs this really better than eleven labs?
- nickolas_t 3y agoCame here to find out myself, still not sure.
- gamegoblin 3y agoJust listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable. Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subscription and just synthesize audio from ebook text. So if I can get a locally running voicebox model and just leave it running on my laptop over night transcribing an audiobook, that's even better.
- bavell 3y ago> So if I can get a locally running voicebox model and just leave it running on my laptop over night transcribing an audiobook, that's even better. This is basically my dream for local AI... locals models trained on my own data/code/styles. Even if they're slow, as long as they work (V/RAM) and are of high enough quality then I'm happy to wait!
- siwakotisaurav 3y agoHave you tried tortoisetts? I believe eleven labs basically forked that and made improvements on voice quality and speed there
- theLiminator 3y agoHow does it compare to Voicebox in quality?
- 3y ago
- chromakode 3y agoI've been working on an open source audio editor which uses Whisper to slice speech. Very exciting to see more capabilities on the horizon!
- theptip 3y ago> As with other powerful new AI innovations, we recognize that this technology brings the potential for misuse and unintended harm. In our paper, we detail how we built a highly effective classifier that can distinguish between authentic speech and audio generated with Voicebox to mitigate these possible future risks. We believe it is important to be open about our work so the research community can build on it and to continue the important conversations we’re having about how to build AI responsibly, which is why we are sharing our approach and results in a research paper Learning from the pushback on releasing LLaMA it seems. I wonder how hard this will be to replicate. (Trained on “60K hours of English audiobooks and 50K hours of multilingual audiobooks in 6 languages for the mono and multilingual setups”, this doesn’t sound intractable.)
- fbdab103 3y agoAssuming 10 hours a piece, 6k books feels a very achievable dataset. Even Librivox claims 18k books (with many duplicates and hugely varying quality levels). If you wanted to get expansive, you could dig into the podcast archives of BBC, NPR, etc which could potentially yield millions of hours. [0] https://librivox.org/ https://librivox.org/
- tudorw 3y agoolder BBC material when the 'standard bbc' voice reigned supreme might make a great training set and could/should be publicly accessible?
- theptip 3y agoFrom the paper: > Model Transformer [Vaswani et al., 2017] with convolutional positional embedding [Baevski et al., 2020] and ALiBi self-attention bias [Press et al., 2021] are used for both the audio and the duration model. ALiBi bias for the flow step xt is set to 0. The audio model has 24 layers, 16 attention heads, 1024/4096 embedding/feed-forward network (FFN) dimension, 330M parameters. We add skip connections connecting symmetric layers (first layer to last layer, second layer to second-to-last layer, etc.) in the style of the UNet architecture. States are concatenated channel-wise and then combined using a linear layer. The duration model has 8 heads, 512/2048 embedding/FFN dimensions, with 8/10 layers for English/multilingual setup (28M/34M parameters in total). All models are trained in FP16.
- poisonarena 3y ago> introduces state of the art AI Model for speech > narrator for the presentation is indian woman with lisp every time
- insickness 3y agoThe narration was bad. I was thinking maybe it was the head developer of the project and they wanted her to narrate or something. Why would you choose that narration?
- stan_kirdey 3y agoseem to be already in the process of reproduction by the community - https://github.com/SpeechifyInc/Meta-voicebox https://github.com/SpeechifyInc/Meta-voicebox
- paul7986 3y agoAnyone know what tool is being used to create AI singing voices/renditions of Mariah Carey singing Whitney Houston to other popular songs? Here's a bunch of results on YouTube and some are really good https://www.youtube.com/results?search_query=mariah+carey+ai+voice+generator https://www.youtube.com/results?search_query=mariah+carey+ai...
- ogsalman 3y agooh okay, so this is what they were recording me for... hehe, fun world we live in.
- jansan 3y agoPretty impressive, but I had the video running in the background and it sounded a bit too sterile for my taste. Also, the narrator sounds a bit like she once gave a foot massage to Marsellus Wallace's bride.
- cvhashim04 3y agoAny available open source APIs?
- treprinum 3y agoSo are they releasing it or not? It's a nice PR statement but "we are not making the Voicebox model or code publicly available at this time". Phantom release?
- fnordpiglet 3y agoThey are holding it closed for the safety of the children and their future profit margins.
- deleted 3y ago[deleted]
- lucidrains 3y agocan an expert in the field comment on whether the results are more or less impressive than Soundstorm?
- fisot 3y agoI've got the same question; found some of the researchers for both projects on twitter and will see if I can get an opinion from one of them. Just waiting on verification to pm them. Will reply here if I hear back unless you have a contact with notifications.
- lucidrains 3y agogot a response here https://github.com/lucidrains/soundstorm-pytorch/discussions/13 https://github.com/lucidrains/soundstorm-pytorch/discussions...
- ftth_finland 3y agoAll I want is a podcast player that automatically cleans up the crappy audio.
- schappim 3y ago> we are not making the Voicebox model or code publicly available at this time Hopefully it’ll do a LLaMA.
- paul_funyun 3y agoLooking forward to being able to automatically generate audiobooks in famous voices. Gilbert Gottfried's "Blood Meridian" will be a hoot.
- deleted 3y ago[deleted]
- jasfi 3y agoThis is like qhwn Google published all those papers about their LLM tech, and ChatGPT just launched something that worked. It will end the same way for their voice tech if they never release it. Someone will release something just as good and take the market.
- chuan_l 3y agoThis is not close to " state of the art " in TTS. The output is clicky , low bitrate and lacks vocal nuance. Its novel for using a " flow matching " approach in its architecture and being suited to cloud - based translation. Have a listen to U Washington , Google " sound storm " instead !
- jmiskovic 3y agoA common courtesy would be to check for naming collisions in voice-generating domain. https://github.com/jmiskovic/voicebox https://github.com/jmiskovic/voicebox If I ever start selling scrapbooks for collecting human faces I'll be returning the favor.