17 ms·
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
- skilled 3y agoRelevant, https://news.ycombinator.com/item?id=37493741 https://news.ycombinator.com/item?id=37493741 https://www.stableaudio.com/ https://www.stableaudio.com/
- kherud 3y agoThank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler than audio generation and the demand far bigger, since not everyone has a professional audio setup ready at all times. [1] https://podcast.adobe.com/ https://podcast.adobe.com/
- earthnail 3y agoWe've been researching an audio denoiser for music that we will present at the AES conference in October. Description page: https://tape.it/denoising https://tape.it/denoising We'll also publish a webapp where you can use the denoiser for free. Mail me if you want beta access to it (email in profile). It won't be open-source though, although the paper will of course be public. It will also only reduce noise, and not reconstruct other aspects of audio quality. However, it can do so on any audio (in particular music), not just speech like Adobe Podcast, and it fully preserves the audio quality. It's designed exactly for the use case you want: to make noisy recordings sound professional.
- white_beach 3y agodenoising seems to fail in the guitar and vocals example
- earthnail 3y agoCan you clarify where it fails? It's designed to remove stationary noise only, and removes it very well in the guitar and vocals example. Generally speaking, if you have other sounds that you don't want in the audio, we don't remove them - it's hard to decide from a musical point of view whether you want a certain sound or not. To give an extreme example: a barking dog probably doesn't belong into a Zoom conference, but it may very well belong into your audio recording. Removing such elements would be a creative decision. The guitar and vocals example has certain clicks in the background that we don't remove - but the stationary noise is gone. Existing professional (and complex) audio restoration tools like iZotope RX don't remove those clicks, either. It's a conservative approach, sure, but in return you can throw any audio at it and it always improves it.
- haywirez 3y agoAre you sure the demo sound files are correct on the website? Couldn't appreciate any glaringly obvious differences between the original and denoised with studio grade headphones here. Or, the originals aren't noisy enough.
- joshspankit 3y agoThere seems to have been a fork in the road: On one side the tech for literal denoising has stagnated a bit. It’s a very hard problem to remove all noise while keeping things like transients. On the other side, AI is being rapidly developed for it’s ability to denoise by recreating the recording, just without the noise.
- earthnail 3y agoIn our denoiser (see other comment), we worked on combining these two forks. That’s how we can mathematically guarantee great audio quality. This combination was non-trivial as training old school DSP denoisers is not easily possible. We’ll describe the math needed in our paper. We hope our publication will help the wider community work not just on denoising but also tasks like automatic mixing.
- whywhywhywhy 3y agoIt’s not open but Nvidia has RTX Voice for free if you have and Nvidia card. Only weird thing it’s designed to be used real time but I’ve had some luck on cleaning up voice recordings replayed back through it via audio routing.
- cosmok 3y agoI have had a lot of success with this: https://ultimatevocalremover.com/ https://ultimatevocalremover.com/ for de-noising
- spdif899 3y agohttps://youtu.be/o-kJ4_CuWzA https://youtu.be/o-kJ4_CuWzA This video from MKBHD's studio channel dives into this topic
- naillo 3y agoI keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.
- Tenoke 3y agoStability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.
- Taek 3y agoBefore stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.
- Tenoke 3y ago>Before stable diffusion, nobody released weights at all. That's not true. There's been a lot of models with weights from every player before Stability. >Without stability, all of AI would still be closed and opaque. Most GANs (the practically spiritual predecessor to diffusion models) for example were available. Huggingface existed and has realistically done more to keep AI open. And again, this specific release we are talking about by Stability is not Open. Stability is great but you are re-writting history and doing it on the release where it makes least sense to do so.
- refulgentis 3y agoNah. Dunno where this is coming from but infamously no AI models were released by big players for years. Rewind 18 months and all you got is GPT-3.0 that no one seems to care about and Disco Diffusion-y type stuff.
- 3y ago
- jncfhnb 3y agoThe bluegrass one is super weird. I can’t identify exactly why.
- benesing 3y agoAlso, the music is not bluegrass as much as it is old-time, a confusion that continually irritates old-time players.
- smat 3y agoYou are right it feels off. The position of the guitar in stereo is all over the place, higher frequency elements appear to come from the left while other parts are more centered.
- stef25 3y agoSame for the death metal
- ewan251 3y agoI think the super weird part is that it's not great? I understand this is most likely very impressive technologically but musically it is disjointed, inconsistent and fake sounding. Most of the "music" examples have weird phrasing and confusing harmonic rhythm. Kudos to stability.ai for achieving this as I am sure it took a lot of effort and this is a huge leap forward in terms of generation of audio by generative AI. However as a musician (BMus and MMus at 2 different conservatoires) I think it's important to say that the job risk being experienced by creative writers will not be extending to musicians... yet.
- jncfhnb 3y agoI feel like music composition is a fundamentally hard task for AI. Music production seems like it should be a lot easier but I haven’t seen that
- viraptor 3y agoFrom what I've seen in the generated tracks so far (this one and others), they're pretty good locally, but just ignore the overall composition. For example any generated blues tracks will have the vague blues feel, but won't keep the 12 bar style. The bluegrass example here doesn't even seem to keep to 4/4 (or is extremely fluid about it...). Maybe one day someone will add a higher level "what's the current section, how far are you into it" inputs to that model to get something better - literally preparing the structure first and then filling it in. That should get much better results for context like "you're playing blues in A with quick change and generating bars 3-4, match the previous bars in style". I mean, chatgpt knows how to plan this out https://chat.openai.com/share/976077c0-138b-4363-8065-3c8eedc42495 https://chat.openai.com/share/976077c0-138b-4363-8065-3c8eed... Painting in that picture should be much easier than generating something freeflowing. Generating a good structure isn't that hard for most styles, because you can literally use the same pattern and do a few random changes that keep the key. (See lots of pop songs using the same 3/4 chord progression)
- jacooper 3y agoEverything looks very convincing apart from the airpane pilot and the sound effects. They sound very weird as if one is hallucinating
- sebzim4500 3y agoThe airplane one just sounds like a foreign language over a bad intercom, I think that could still be useful for some stuff.
- xpe 3y agoPerhaps because generating good white noise requires randomness without autocorrelation or detectable patterns.
- wiz21c 3y agoThe airplane is super convincing as an encrypted Empire communication :-) (see Star Wars episode 5 IIRC)
- naillo 3y agoThis is gonna be great to finetune on. There's only so many boards of canada/aphex twin songs out there but I wish there were more and this will let us generate more.
- 52-6F-62 3y agoThis is not the way.
- naillo 3y agoWhy not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.
- wokwokwok 3y agoCome on, be creative and make something new instead of copying someone else. It’s just kind of lame imo. “Mostly” private use? Mmm. :thumbs down emoji:
- naillo 3y agoI meant private use and maybe share with a few friends. I actually agree with you that we probably shouldn't finetune on great artists and try to sell the output without modification or added creativity. Private or close friends sharing is fun and life enriching and inspiring though in my eyes.
- 52-6F-62 3y agoBoards of Canada came to their sound because in their youth, the brothers had to move to Canada for a time. Even though it was only a couple of years their experience made an indelible mark on them—particularly school days watching old National Film Board of Canada tapes on worn VCR heads. When they moved back to Scotland and started their music they started incorporating both the machinery and the sounds from the tapes in their compositions. And they could play their compositions live. It was quite the rig. It’s not just entertainment. It’s communicating a very specific feeling and perspective. Keep learning and create, don’t be satisfied with just copying. The biggest difference here is in the doing. You have to grow into one mode over time and energy spent, the other is immediate gratification with minimal personal energy. Everything valuable comes during the course of that process of growing and committing energy. And it’s so good. Don’t deny yourself.
- 2Gkashmiri 3y agoSo.... Wait for llama for audio and train your own voice to having to call you friends by text and the software instead of actually saying the words? This is going to be nice for authentication, proving to a third party that you are yourself
- systoll 3y agoJust wait for iOS17 https://youtu.be/oMt02DNbQlk https://youtu.be/oMt02DNbQlk
- hubraumhugo 3y agoNow imagine Spotify using this to generate individual earworms for everybody based on their personal tastes (likes, playlists). Yes, AI is partly hype, but had someone told me this even two years ago, I wouldn't have believed it.
- joshspankit 3y agoThis is why it’s vital that AI is openly available. Imagine a world where Spotify is the only company that can do that, and they use it to make sure they never pay royalties again.
- bee_rider 3y agoHow is Spotify for finding new music based on your tastes? I’ve only used Amazon and Pandora; Amazon is quite poor, Pandora is pretty good. I suspect (although, without proof) that if a service can’t suggest new music, it will have trouble generating new music as well. Anyway, I very much would rather run this sort of thing locally. You could just manually set your taste profile. Plus, music can be quite personal, imagine you start listening to too much music inspired by The Cure and suddenly Amazon starts advertising black makeup and antidepressants or something like that, it would be too disconcerting.
- magicalhippo 3y ago> How is Spotify for finding new music based on your tastes? I haven't tried any alternatives really, but so far for me I'd say decent. I put on an album, and once over it'll play similarish stuff. If I don't like a song I'll skip it and it seems to incorporate that feedback. Only thing is that it doesn't seem to be too adventurous and it adheres rather strictly to the local context. Meaning, if I played a stoner rock track, it'll continue suggesting stoner rock and not much else, even though I have quite varied music favorited in my library. Overall though I've found a lot of new bands I enjoy that way so, positive experience for me. edit: as an example, here are the two most recent ones it suggested where I ended up buying the albums on Bandcamp. Both have quite few monthly listeners (1-2k), so not what I'd call mainstream. SUIR https://open.spotify.com/artist/6zOeQ2hyNfqi9UMHtyTSlF https://open.spotify.com/artist/6zOeQ2hyNfqi9UMHtyTSlF Mount Hush https://open.spotify.com/artist/13clfeXxTPsDsqzSlLIBZJ https://open.spotify.com/artist/13clfeXxTPsDsqzSlLIBZJ
- gyumjibashyan 3y agoThis is crazy tech!
- colesantiago 3y agoThis is yet another amazing release from Stability AI. Will be adding this to my SaaS side grift and introduce generated music you can listen to while you're chatting with your PDFs. Can't wait for the next one.
- jasbur 3y agoIt's interesting that the Death Metal was the hardest to reproduce. I conclude that it's the most fundamentally human of all genres.
- Ninovdmark 3y agoThe sound sample seemed to fit the 'vibe', but lacked any discernible definition. Could it be that it's too sonically dense to easily reproduce? Perhaps this could be improved with a more tailored training set.
- awestroke 3y agoI think it was just the genre that was the least represented in the training data
- chankstein38 3y agoIt sounds more like break core haha
- dontreact 3y agoWell... they hardly tried all genres :) It sounds like it can't handle lyrics or semantics that well so I suspect any genre where the lyricism is important would also be quite mushy and recognizably AI
- Jeff_Brown 3y agoThe Beatles seemed to be the hardest music for JukeBox to emulate.
- emperorcj 3y agodadabots here: haven't gotten good death metal with it. problem is there's not really much in the dataset.
- emperorcj 3y ago(context: i make ai death metal & also i worked on stable audio). 100% it was a dataset problem. Diffusion models still work well when you train them on death metal: https://www.youtube.com/watch?v=rlsRMQzD_6Q https://www.youtube.com/watch?v=rlsRMQzD_6Q
- coldcode 3y agoIt's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, though I doubt you could run the AI generator inside a game, you could generate it as a asset. People like Hans Zimmer and Ludwig Göransson have nothing to worry about. Singing would an interesting experiment, but I don't see that here.
- Maschinesky 3y agoI'm surprised that you even leap to people like Hans Zimmer and others. The people we need to worry about are aallll of the people earning a living for everything else like background music for Indi games, Ambient music etc.
- riskable 3y agoI thought the, "epic trailer music intense tribal percussion and brass" was pretty good. Rather, good enough for something like a video game where the game engine is dynamically generating music based on the present situation in the game. I could easily find that music entertaining if it started playing the moment my character triggered a trap and suddenly, "the floor is lava" or my character enters a scene with the quest of winning over one of the love interests =)
- zone411 3y agoYes, while working on my AI Melodies Assistant project, it quickly became clear that generating a pleasant but boring music isn't too difficult. To create a catchy tune, an element of surprise is essential. In the end, I was able to use it as an assistant to compose 60 melodies that I'm happy with (https://www.melodies.ai/ https://www.melodies.ai/)
- PatronBernard 3y agoWhat I haven't seen done well by these generative AI's so far is structure (having a chorus, a verse, a bridge, ...) and harmonic movement/progressions (except for maybe a V-I or I - VI - ii - V) And those two be things are exactly what makes a song interesting and non-repetitive.
- Cloudef 3y agoIs the extreme metal music lacking from the training set? Why do the extreme metal examples always sound horrible?
- rafaelero 3y agoBecause metal sounds horrible.
- Jeff_Brown 3y agoMetal is especially hard to mix in a way that keeps the voices distinct and clear. Maybe the training catalog includes a lot of low-budget metal.
- iandanforth 3y agoThe solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to output full images, but either layers or brush strokes and tool pallet usage. Same for organizations that have all the component tracks of music to work with.
- Jeff_Brown 3y agoThat raises an interesting difference between cleaning AI-generated sound and cleaning ordinary recordings. In an ordinary recording, there is an objective reality to discover -- a certain collection of voices was summed to create a signal. With (most? the best?) existing AI audio generation, the waveform is created from whole cloth, and extracting voices from it is an act of creation, not just discovery. I've come across AI-generated music that outputs something like MIDI and controls synthesizers. Its audio quality was crystal-clear, but the music was boring. That's not to say the approach is a dead-end, of course -- and indeed, as a musician, the idea of that kind of output is exciting. But getting good data to train something that outputs separate MIDI-ish voices seems much harder than getting raw audio signals.
- fnordpiglet 3y agoGenerative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.
- jskherman 3y agoI believe Spotify's Basic Pitch[0] is already some work towards building something like this. [0]: https://basicpitch.spotify.com/about https://basicpitch.spotify.com/about
- stared 3y agoI would love to use it for background music when I am working. I have specific tastes that depend on the task, mood, energy level, and ambiance.
- k12sosse 3y agoIf you're not attuned to Cryo Chamber (label), check them out. Maybe not fitting all use-cases, but a strong and deep catalogue.
- TheAceOfHearts 3y agoDoes this model support / "understand" concepts of spatial audio? For example, something like "an alarm moving around you in a circle". When AudioGen was announced this was my first question, but from what I've been able to test the model just ignores spatial audio prompts. Unfortunately I haven't been able to find any discussion or interest in online discussion about the importance / significance of spatial audio. Why not?
- Jerrrry 3y agoDolby wouldn't appreciate it.
- cheald 3y agoMy guess is that it's not a very interesting problem because it's not particularly difficult to add spatial dimensions to arbitrary audio - after all, it is already commonly done in video games. All you have to do is manipulate the multichannel outputs with an understanding of the spatial positioning of each channel's speaker location relative to the listener and some basic trig.
- PcChip 3y agoit's funny how they're all very impressive except the death metal
- Jeff_Brown 3y agoI still consider OpenAI's JukeBox (now at least 2 years old!) far and away the most creative music AI. But the combination of coherence, sound quality and creativity of this model is (to my knowledge) easily best in class.
- 4RealFreedom 3y agoThe sound quality of Jukebox is muddled. There are many inconsistencies. The loudness of vocals and the quality of instruments really stand out and not in a good way. Hard to talk about creativity because it's so subjective but I've found it lacking in all AI music including JukeBox. Don't get me wrong - this tech is amazing.
- Jeff_Brown 3y agoIt's mushy and inconsistent, absolutely. But it also comes up with wild yet coherent changes that I've seen from nothing else. At this comment I listed a few instances: https://news.ycombinator.com/item?id=37499067 https://news.ycombinator.com/item?id=37499067
- dylan604 3y agoJust like all examples of generative "AI" I've seen, there's always some bit of uncanny valley vibe present. In the audio examples, there's always this weird distortion like a really poorly compression sources were used as training data. The sounds are muddled together, and rarely do I hear clean musical voices. It's just a smear of sounds coming together that our brains try really hard to say "oh, that's a _____" situation. While the samples in the TFA are probably the closest I've heard to date, the issue is still present. I guess the thing that strikes me so odd about the generative thing is all of the press releases on people presenting things like it's a final product, yet it's clearly pre-release beta at best but more likely alpha versions of code in the results in quality. If a non-AI product released something that was so clearly not finished, it would be panned to no end for not working.
- deleted 3y ago[deleted]
- skybrian 3y agoAs an amateur musician, I’d be more interested in these tools if, along with the text description, they took as input a melody or chord progression or performance data. Maybe ABC notation or a MIDI track? Anyone doing that? Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. There could be a new generation of audio tools that let you generate placeholders or better for everything.
- l33tman 3y agoThe analogue from stable diffusion would be ControlNet, where you can train a superimposed model on auxiliary data, this should be possible to do with chords for example, just like you can do with human poses, 3D depth maps etc in stable diffusion using controlnet
- hoosieree 3y agoHumans take a long time to get good at art; in the meantime they still have to eat. So they compete with generative AI for a fixed number of jobs. The AI is cheaper and faster. Humans stop training to become artists. Without new training data, the generative AI models stagnate. Progress in art stops globally, forever. But for a brief glorious moment, we were able to say "huh, that's not bad".
- phone8675309 3y agoThis is by design - the capital that backs modern art isn’t doing it for love of the art but for money. For fine art, it’s a way for them to launder money and keep it out of bank accounts where it can be seized trivially. For mass art, it’s about selling to enough rubes to make a profit. Neither are impacted by a stagnation in art. If anything, they’re aided by it - suddenly the art you bought to launder money retains its value because it’s no longer the flavor of the week with the arts crowd.
- deleted 3y ago[deleted]
- randcraw 3y agoI wonder if it makes sense to generate a combo of instruments rather than individual voices and then combine those with an arranger DNN. I would think it'd be much easier to capture each instrument's transients and dynamics that way, much less allow more subtlety in how they combine, like allowing the lead voice to shift among instruments, or even let the listener choose how each voice expresses stylistically and how they should combine. Trying to do all of that in a single DNN, much less parameterize it useably seems overly ambitious (or will be of more limited value ultimately).
- shon 3y agoEd Newton-Rex, VP of Audio at Stability, is speaking about how this was built at The AI Conference in 2 weeks. https://aiconference.com https://aiconference.com
- cesaref 3y agoThe death metal example reminded me of the continuous streaming death metal here: https://www.youtube.com/watch?v=MwtVkPKx3RA https://www.youtube.com/watch?v=MwtVkPKx3RA
- _sys49152 3y agogamechanging stuff for sample based rap producers. havent been able to log-in yet but i think a good benchmark to start off with is to see if it can replicate the 'al green' sound from the early 70s - very distinct sounding production - drumless and instrumental. you dont need 45 or 90 straight seconds of a coherent song rendered. just need to dip in the 45 sec clip and cut out 4 seconds here, another 4 there. reroll those cuts through stable audio, keep rolling, keep rolling. cut up and get a pile of clips together. arrange, layer, voila - you saved money on paying royalties for sampling. the lofi melodic sample on the stability page was passable. thought the bluegrass one sounded great actually. imagine being able to program bluegrass like rap. edit: oof. fully trained on a licensed commercial dataset from AudioSparx. muzak in, muzak out.
- stainablesteel 3y agoi want something that can take in a song and transform it into a different genre
- smrtinsert 3y agoAs a hobby musician, why don't they start with instrument samples? Every sampler user out there would love a button press generate sample on the fly as a plugin. It would blow away gigs and gigs of ridiculous duplicative or near duplicate samples.
- londons_explore 3y agoSeems nobody has really cracked vocals with songs...
- junon 3y agoThis is the most impressive of this genre of SD audio so far, by a long shot. Really impressive!
- moonchrome 3y agoIt's funny how some of those examples give me this creepy uncanny valley feel for music (the lowfi hip hop example) - I've never experienced it this way before. It's sort of reminds me of the audio effects they use to indicate that you're incapacitated and things start distorting in a weird way. Entertaining !
- chandhoo 3y agoAnother alternative: http://www.word.band http://www.word.band Can produce longer content, and more genres and range of music. Isn't 48khz though.