22 ms·
Audiobox: Meta's new foundation research model for audio generation
- youssefabdelm 3y agoHave weights been released? Edit: nvm, seems not from this line "In the coming weeks, we will be opening up the application here, along with an interactive demo that will showcase Audiobox’s capabilities."
- nuz 3y agoVR is gonna get wild in like 5 years if they keep this up
- spaceman_2020 3y agoHave you seen some of the demos people have been building with Unreal 5.3? Insane stuff. In a decade, this stuff will be hard to tell from reality.
- SheinhardtWigCo 3y agoSounds cool - got any specific ones to share?
- dvh 3y agohttps://youtu.be/A7tp4eg0ax8 https://youtu.be/A7tp4eg0ax8
- jl6 3y agoThe geometry and lighting is amazing but I couldn’t detect any animation, like gentle motion due to wind.
- flyaway123 3y agoSome motion due to wind, in another video: https://youtu.be/_B9hkn6wgNA?feature=shared&t=24 https://youtu.be/_B9hkn6wgNA?feature=shared&t=24
- pants2 3y agoThat's really good, though I also want to point out some of the amazing graphics that modders accomplished in Crysis (2008): https://youtu.be/3w6COXBfIY4 https://youtu.be/3w6COXBfIY4
- jay-barronville 3y agoIncredible.
- Racing0461 3y agoThe new avatar (blue skin not arrow) game looks like this demo with some characters tossed in to control.
- prakhar897 3y agohttps://www.youtube.com/watch?v=IK76q13Aqt0 https://www.youtube.com/watch?v=IK76q13Aqt0
- nerdix 3y agoThis is why I'm high on the metaverse long term. In ten years, there will be a $500 (or whatever the 2033 inflation adjusted value is) VR headset that blows the Apple Vision Pro out of the water in terms of optics, will run a highly optimized version of the lastest revision of Llama locally (and it will be much better than anything we currently have today), come with wifi 8 (so it will have multigigabit per second real word performance), capable of rendering graphics that look much more realistic than Unreal Engine 5 (with high frame rates due to AI upscaling and frame generation). There will be people that will spend almost every waking hour with one of those things attached to their face if they can also make this device lightweight and comfortable
- holoduke 3y agoIn 2040 we will have lenses with 32k displays and gpus with 10 trillion transistors and 1 petabyte of memory. Its hard to predict what you can do with that. The real world would be empty by then.
- TaylorAlexander 3y agoThe real world is much richer than a bunch of stuff just being projected in to your eyes. You can’t climb a tree in VR.
- kthartic 3y agoMaybe for some. But for those living in unfortunate circumstances (low socioeconomic status, small apartment, poor future prospects, etc) I bet a VR world of paradise and social connection is far more inviting.
- doublerabbit 3y agoAnd yet I'll still be waiting five weeks to just download the world because I'm stuck on 2Mb/s ADSL.
- morbusfonticuli 3y ago
- justapassenger 3y ago[dead]
- 9dev 3y agoIf I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, and motivations. And those virtual conversation participants can talk back to you, react to your words and actions in a believable, fully immersive manner. That's a dream come true for every gamer on the face of this earth, I believe. The more rational voices in my mind, though, become more and more afraid of a world where the only thing you can trust is people sitting right in front of you. That makes the world of information pretty small again.
- logicchains 3y ago>If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, and motivations. And those virtual conversation participants can talk back to you, react to your words and actions in a believable, fully immersive manner. That's a dream come true for every gamer on the face of this earth, I believe. It's basically Dungeons and Dragons with an AI dungeon master who can generate video in realtime. Which would be awesome, but like Dungeons and Dragons it wouldn't be easy to keep the player on track.
- 9dev 3y agoI'd imagine everyone else has an agenda, a schedule and ordinary life which is pre-written, so the game actually wants to tell an interesting story, but as a player, you can choose to listen or just do all shenanigans in the context of the game you can come up with. If you want to be a farmer instead of defeating the dark overlord, so be it -- until the world ends because the overload has achieved his goals with no resistance, or maybe someone else becomes a hero instead. ...Gosh, the more I think about it the more awesome it gets. Edit: just one more! Imagine actually having to complete quests in a given amount of time, because the rest of the world continues to revolve. People being mad at you because you left their children to die in the dungeon after arriving a day too late, because you were busy running an errand for someone else.
- nathanfig 3y agoMulti input? Infilling? First generative audio model I've seen that starts to close the gap with image models.
- novolunt 3y agoI think that for artificial intelligence to become like humans, it should be treated under the same conditions as natural humans. It should be able to see the surrounding environment, listen to the surrounding sounds, smell the surrounding smells, and taste the surrounding food. It should be given Parents and relatives should be given their own partners and their own country. In this way, the artificial intelligence trained in the environment will naturally be more like human beings and have their own emotions.
- kevindamm 3y agoWhat you're looking for is embodiment, and actually this was explicitly left aside in a recent paper that attempts to give measurement criteria for AGI[0]. But I agree with you that the entire lived sensation is critical to approaching any objective involving alignment. [0] https://arxiv.org/abs/2311.02462 https://arxiv.org/abs/2311.02462
- empath75 3y agoWhat if the "environment" for an AI is just "the internet". A long time ago, there was a great story in a Shadowrun supplement of all things about a hacker that got trained to teach an "ai" how to break into computers. It was basically at a child's level, emotionally, and the only world it's ever known was "the matrix" (yes, really -- and written almost a decade before The Matrix came out). Eventually it turns out that it's not an ai, but a corporation was stealing kids and sticking them in Virtual Reality at birth to train a team of super hackers.
- maroonblazer 3y ago> We’re inviting researchers and institutions who have been previously involved in speech research, and who want to pursue responsibility and safety research on the latest Audiobox models, to apply. It's not clear as to what the expected outcome of this 'responsibility and safety research' effort is. Is the idea to nerf the tech such that it can't be used for morally/ethically nefarious purposes? If so, then is the "speech research" community the group best fit to do that work?
- david_draco 3y ago[flagged]
- deleted 3y ago[deleted]
- varunytoons 3y agoThis a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio https://github.com/QwenLM/Qwen-Audio Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understand is that the model weights have to be swapped or altered per se to be able to commercially reuse this. Correct me if I'm wrong.
- two_in_one 3y agoThanks for the link. License is clear: Researchers and developers are free to use the codes and model weights of both Qwen-Audio and Qwen-Audio-Chat. We also allow their commercial use. and important, if you have more than 100m active users:) 4. Restrictions If you are commercially using the Materials, and your product or service has more than 100 million monthly active users, You shall request a license from Us So, looks like it's absolutely fine to use, except for IT behemoth. As for how to use, API, I think. Interesting applications are possible. Like interactive mobile robots. Assistants for people with disabilities, both software and wearable. Interesting times... this will be called AI revolution probably. It's already not a joke, after several ups and downs.
- Zuiii 3y ago> Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio https://github.com/QwenLM/Qwen-Audio Openly distributed perhaps, but definitely not open source. The license appears as closed as meta (research-only with some leeway for other uses.) Do you have any truely open-source general audio generation models yet? I know about StyleTTS2, which is open source (MIT) and uncensored, but that model focuses on speech generation only. Having an non proprietary model like audiobox or Qwen-Audio would be really nice.
- TaylorAlexander 3y agoNormally I want basically everything to be open source. But as soon as I saw that audio restyling demo, I began to feel concerned they may be open sourcing this. The model can take a sample speakers voice, new text to speak, and also a description of a new location (like a cathedral with many echoes, or other background noises) and produce a convincing new audio sample. This technology will present serious challenges for the verification of covertly recorded audio. It will of course ultimately become widespread but I’m not inherently bothered by the idea of slowing down its release. Giving researchers extra time to examine possible detection techniques seems helpful to me.
- mmaunder 3y agoI think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be eating leftovers and paying a fortune for them, wondering why you can’t keep up - unless we enable a vibrant open source ecosystem, and force big tech to release models into that ecosystem. Support open source models by celebrating their release and pressuring companies to release them, and oppose closed source AI or face a very bleak future for you and your descendants. You may be having fun with “Open” AI’s API today, but you’re supporting and celebrating the collapse of society into megacap AI elites and a majority paying for metered access to old technology.
- s3p 3y agoI mean sure.. but imagine if this were open sourced as is. This is new tech that has barely had time to mature. The possibilities for abuse are endless. I for one am happy that this model isn't being open sourced. This is an excellent way for people to generate all kinds of disturbing and fake audio clips.
- mmaunder 3y agoSame logic could be applied to Linux by Microsoft in the 90s. In fact, the “It’s for your safety” has been applied to some of the worst things humans have perpetrated including apartheid (which I lived through) and the holocaust. And it’s always those that claim to keep us safe doing the worst. And it continues with perceived dangers providing pretext and moral authority to do bad things.
- beebeepka 3y ago"If we were to release our kettle and chickens, they would go extinct within a few years." "I just hold on to all the money, 'cause bitches can't be trusted with it. We pool all the kissing money together, see? But if you wanna buy anything, you just talk to the bottom bitch, and then the bottom bitch talks to me. Do you know what I am saying?"
- petarb 3y agoWhat’s the best way to try these models out? Does Meta usually provide a web interface for them or do you have to download and run locally?
- noiseinvacuum 3y agohttps://audiobox.metademolab.com/ https://audiobox.metademolab.com/
- holoduke 3y agoHow long before someone manages to clone him/herself online and apply for relatively simple gig work. And duplicates than 1000 times. Making millions with simple work. Its almost possible i think. Clone your voice with this, clone your looks by wiring comfyui like sd nodes to your webcam. Everything instructed/orchestrated by some AI agent controlled by chatgpt. Some wiring logic is what you need to make. The only thing you as the real person have to do is answering some critical decisions which are send by a push notification.
- TaylorAlexander 3y agoBy the time a machine learning model can replace 1000 workers they’ll just stop hiring real workers. What remains will be tasks which can’t be automated.
- nuz 3y agoOr isn't allowed to by law, e.g. security etc.
- a_wild_dandan 3y agoRegarding their "responsible" model, Meta's engineers aren't stupid. They know that: 1. No TTS audio output is tamper-proof. Their "safeguards" will be busted, and quickly. Whether via a small adversarial NN, some basic DSP, or just...holding a cheap recorder near your speakers, maintaining audio file provenance has no chance. 2. Impersonations have vexed humanity since the invention of vocal cords. Insofar as it's soluble, it's been solved -- authenticity is determined by a fluid mixture of context, trustworthiness, and the authority of involved parties & institutions. Always has been. Always will be. If I could drill one idea deep into every tech evangelist's head, it'd be: The solution to every problem isn't automatically "more technology." But hammers see only nails, so the vicious cycle continues, and society deals with the consequences (e.g. cryptobros decentralizing money...by slowly reinventing banks, but with more fraud). 3. This secret audio ID "feature" is probably harmful. It adds needless complexity. At best it exacerbates a false sense of safety because impersonation is trivial. Bad guys can emulate it on authentic recordings to discredit them as "fake." Nobody who'd actually benefit from such safeguards will respect them. News says this audio that affirms my confirmation bias is fake? Nah, the news is fake. Meta knows all of this. Optimistically, I hope it's just lip service to concern fetishists; plausible deniability for the knife manufacturer when a bad guy uses one. Pessimistically, it might be pretext for an about-face on their OSS commitment. "Oops, researchers trivially broke our safeguards. Shucks. That's scary. Guess we'll build a moat instead of an OSS community. Think of the children or terminator or whatever works these days" I suppose we'll see.
- kkzz99 3y agoThe speed of progress is just incredible. I've been using all kinds of different TTS engines for years now and the rapid pace of advancement is awesome. I usually generate all my audiobooks from ebooks and articles and the quality and stability (think artifacts) has gone up so much in the last few months.