13 ms·
Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Mea
by haykmartiros 4y ago
Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day!
Meanwhile, please read our about page http://riffusion.com/about http://riffusion.com/about
It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself
This has been our hobby project for the past few months. Seeing the incredible results of stable diffusion, we were curious if we could fine tune the model to output spectrograms and then convert to audio clips. The answer to that was a resounding yes, and we became addicted to generating music from text prompts. There are existing works for generating audio or MIDI from text, but none as simple or general as fine tuning the image-based model. Taking it a step further, we made an interactive experience for generating looping audio from text prompts in real time. To do this we built a web app where you type in prompts like a jukebox, and audio clips are generated on the fly. To make the audio loop and transition smoothly, we implemented a pipeline that does img2img conditioning combined with latent space interpolation.
- jtode 4y agoAs one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.
- spoiler 4y agoAs a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a familiar word in a foreign language; the last one can happen more often though). So, I think these tools will be used to do background work (ie for visuals maybe help with background tasks in CGI or faraway textures in games). I know less about audio, but I assume it could maybe help a DJ create a transition between two segments they want to combine, as opposed make the whole composition for them, but idk if that example makes sense. Now, onto a more human point: I think that people often listen to music because it means something to them. Similar for people who appreciate visual art. I also love interactive and light art, and I love talking to other artists at light festivals who make them because of the stories and journeys behind their art too. Humans and art are a package deal, IMO. Edit: typos and to add: Also, I think prompt authorship is an art unto itself. I'm amazed what people can craft with it, but I'm more impressed by the craft itself than the outputs. Don't get me wrong, the outputs are darn cool, but not if you look closer. And it's impossible to look beneath the surface altogether, as there is nothing in the output but the pixels.
- api 4y agoIn general all this stuff is chopping the bottom off the market. AI art, code, writing, music, etc. can all generate passable "filler" content, which will decimate all human employment generating same. I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market. That being said the bottom of the market is how a lot of artists make their living, so this is going to deeply impact all forms of art as a profession. It might soon impact programming too because while GPT-type systems can't do advanced high level reasoning they will chop the bottom off the market and create a glut of employees that will drive wages down across the board. Basic income or revolution. That's going to be our choice.
- gummydogg 4y agoThe top of the market started at the bottom. Entry level is requiring higher and higher skills and capabilities.
- astrange 4y agoThe only thing that affects whether you have a job is the Federal Reserve, not how good productivity tools are. You always have comparative advantage vs an AI, so you always have the qualifications for an entry level job. There will never be a revolution and there's no such thing as late capitalism. Well, not if the Fed does their job.
- halkony 4y agoI see a lot of AI naysayers neglecting the comparative advantage part. If AI completely eliminates low skill art labour from the job pool, it's not like those affected by it are gonna disintegrate, riot, and restructure society. They have the choice of filling an art niche an AI can't or they can spend that time learning other, more in-demand skills. This also ignores that fact that some companies would rather reallocate you to more profitable projects even if your art skills don't change. Selling a product with relative value like a painting or a sculpture will always be an uphill battle. Now that there's more competition from AI, it just gives artists/businesses incentive to find what people want that an AI can't deliver. Worst case scenario, employment rates in this sector are rough while the market recalibrates. Interested to see how these technologies develop.
- gravelc 4y agoI love simple generative approaches to get ideas, and go from there. This seems like an extension of that (well, it's what I'm going to try - sample the output, make stems, pull MIDI etc). Will make the creative process more interesting for me, not less. Having said that, it's not my job, and I can see where the issues lay there.
- godelski 4y agoWhy do you think this will kill your job? To me this looks like an extension of the hip-hop genre.
- jtode 4y agoI am an active musician, but I don't actually make money at it, I was mostly joking, but: I believe that we are (one determined smart person + six months) away from bots on Youtube and other streaming platforms that generate endless "new" music that follows those simple formulas (beat, bass, sample = loop, several loops connected up = song) 24/7. Raves that have no human DJs and never stop. Good? Bad? Not for me to say, really, the most I make at it is a couple hundred bucks for two nights in a bar playing FM radio hits, and there's lots of people younger than me who like that music, so obviously I'm doing it for different reasons and I don't anticipate losing access to as many bar gigs as I want for the rest of my life. But certain genres are very tolerant of low-effort music, and I think the people who are monetizing low-effort music are gonna lose their income streams. I do different things than those people, but I still consider them compatriots, even if I don't care for their art.
- WheelsAtLarge 4y agoThese are tools. Don't think of them as replacements, they aren't. But as tools that will help us be creative. As smart as these apps seem, they will still need a human to decide where and how to use them. They won't replace us but we need to adapt to a new reality.
- coldtea 4y ago>As smart as these apps seem, they will still need a human to decide where and how to use them. They won't replace us Well, they will, if the AI plus 1 human deciding "where and how to use them" can replace producers and musicians playing...
- atleta 4y agoI hear this a lot (in relation to various jobs) and I still don't get it. Yes, it is a tool. Yes, if it can, it will replace humans. That's the whole point. For some reason people tend to think that these tools/AI/ML systems will never be good enough to do their job (or a specific job). This argument can take different forms, sometimes stating that it will just do the boring part of the work (e.g. with programming) or that it will still need human creativity (maybe, but not necessarily and that's not the point) or that it will just replace low level, unskilled or mediocre professionals. And somehow everyone thinks they are not mediocre (i.e. average). But even these assumptions are unfounded. Why would anyone think that these systems will top out below their skill levels? Why would anyone think that they can't become superhuman? They did in chess, go, I think poker too. Not to mention protein folding. And without much of a hitch between mediocre/good enough and superhuman. Because that difference is just interesting for us, but doesn't necessarily mean that there is huge step, that the system needs to undergo serious development and that it would take a long time. (Like decades or so.) People thought that was the case when AlphaGo beat Fan Hui saying that Lee Sedol was a completely different level. Which, of course, he is. Still, it just took DeepMind half a year to improve alphago to that level. So yeah, you can be pretty sure that if this track (no pun intended), if this solution is good enough then it will quickly evolve into something that will replace some music creators.
- cmsonger 4y agoKurt Vonnegut and Player Piano has a message for you.
- greenhearth 4y agoIsn't this just a sampler with extra steps?
- trynewideas 4y agoI can't think of a genre that would embrace it faster. The pay-for-knock-off rap beat market will feel more pressure from this kind of tool, especially as loop-oriented as it already is.
- CapsAdmin 4y agoWhen you say fine tuned do you mean fine tuned on an existing stable diffusion checkpoint? If so which? It would be very interesting to see what the stable diffusion community that is using automatic1111 version would do with this if it were made into an extension.
- haykmartiros 4y agoYes from https://huggingface.co/runwayml/stable-diffusion-v1-5 https://huggingface.co/runwayml/stable-diffusion-v1-5. Our checkpoint works with automatic1111, and if you'd like to make an extension to decode to audio, it should be pretty straightforward: https://github.com/hmartiro/riffusion-inference/blob/main/riffusion/audio.py https://github.com/hmartiro/riffusion-inference/blob/main/ri...
- Metus 4y agoCan you run this on any hardware already capable of running SD 1.5? I am downloading the model right now, might play with this later. Guessing at the speed with which AI is developing these days someone is going to have the extension up in two hours at most.
- ronsor 4y agoI bet the AUTOMATIC1111 web UI music plugin drops within 48 hours.
- enlyth 4y agoI have made a basic version here: https://github.com/enlyth/sd-webui-riffusion https://github.com/enlyth/sd-webui-riffusion
- haykmartiros 4y agoYes! Although to have real time playback with our defaults you need to be able to generate 40 steps at 512x512 in under 5 seconds.
- ozten 4y agoAmazing work! Did you use CLIP or something like that to train genre + mel-spectrogram? What datasets did you use?
- teruakohatu 4y agoI was very surprised this was not mentioned.
- deleted 4y ago[deleted]
- newobj 4y agoAll the AI music I’ve heard so far has a really unpleasant resonant quality to it. Why is that? Can it be removed?
- hyperbovine 4y agoPresumably for similar reasons that the vast majority of AI generated art and text is off-puttingly hideous or bland. For every stunning example that gets passed around the internet, thousands of others sucked. Generating art that is aesthetically pleasing to humans seems like the Mt. Everest of AI challenges to me.
- andybak 4y agoI think your comment is off-topic to the post you are replyng to. That wasn't asking about the general aesthetic quality - more about a specific audio artifact. > For every stunning example that gets passed around the internet, thousands of others sucked. From personal experience this is simply untrue. I don't want to debate it because you seem to have strong feelings about the topic.
- hyperbovine 4y agoEven if you remove the artifact, the exact same comment applies. It generates a somewhat less interesting version of elevator music. This is not to crap on what they did. As I said, they underlying problem is extremely difficult and nobody has managed to solve it. I don't feel strongly about this topic at all.
- indigochill 4y ago> It generates a somewhat less interesting version of elevator music. This iteration does, but that's an artifact of how it's being generated: small spectograms that mutate without emotional direction (by which I mean we expect things like chord changes and intervals in melodies that we associate with emotional expressions - elevator music also stays in the neutral zone by design). I expect with some further work, someone could add a layer on top of this that could translate emotional expressions into harmonic and melodic direction for the spectrogram generator. But maybe that would also require more training to get the spectrogram generator to reliably produce results that followed those directions?
- lisper 4y agoWow, I am blown away. Some of these clips are really good! I love the Arabic Gospel one. John and George would have loved this so much. And the fact that you can make things that sound good by going through visual space feels to me like the discovery of a Deep Truth, one that goes beyond even the Fourier transform because it somehow connects the aesthetics of the two domains.
- tbalsam 4y agoI can simultaneously burst a bubble and provide fuel for more -- the alignment of the intrinsic manifolds of different domains has been an interesting research topic for zero shot research for a few years. I remember seeing at CVPR 2018 the first zero shot...classifier, I think? That if I recall correctly trained in two domains that were automatically basically aligned with each other enough to provide very good zero shot accuracy. Calling it a Deep Truth might be a bit of an emotional marketing spin but the concept is very exciting nonetheless I believe.
- lisper 4y agoMy characterization of it as a Deep Truth might just be a reflection of my ignorance of the current state of the art in AI. But it's still pretty frickin' cool nonetheless.
- logicallee 4y agoAlright so this is a pretty amazing new development. I want to tell you something about what the state of the art is in AI. When you wrote that it is a deep truth it was before I actually listened to the pieces. I had just read the descriptions. At the time, I thought that you were probably right because I was thinking that music is only pleasing because of the structure of our brains it's not like vision where originally we are interpreting the world and that's where art comes from. Music is purely sort of abstract or artistic. However, after I listened to the pieces, I realised that they really sound exactly like the instruments that are making the physical noises. For example it really sounds exactly like a physical piano. So I don't know about a deep truth, but it does seem that there is a physical sense that the music represents which it can successfully mimic using this essentially image generating capability. One thing about all of these amazing AI development, is that I still make some long comments by dictating to Google. When it first got to the point that it was able to catch almost everything that I was saying I was absolutely blown away. However, it's really not that good at taking dictation, and I have to go back and replace each and every individual comma and period with the corresponding punctuation mark. Seeing such an amazing developments happening month after month year after year it makes me feel like we are really approaching what some people have called the singularity. When I read about a net positive fusion being announced my first instinct was to think oh of course it's now that that ChatGPT is available of course announcing a major fusion breakthrough would happen within days to weeks it just makes perfect sense that AI's can solve problems that have have confounded scientists for decades. To see just how far we still have to go take a look at how this comment read before I manually corrected it to what I had actually said. -- [I copied and pasted the below to the above and then corrected it. Below is the original version. This is how I dictate to Google sometimes, on Android. Normally I would have further edited the above but in this case I wanted to show how far basic things like dictation still have to go. By the way I dictated in a completely quiet room. I can't wait for more advanced AI like ChatGPT to take my dictation.] Alright so this is a pretty amazing our new development period I want to tell you something about out why the state of the heart is is in a i period when you wrote that it is a deep truth it was before I actually listen to The Pieces, I have just read the descriptions period at the time, I thought that you were probably right because I was thinking that music is only pleasing because of the structure of our brains it's not like vision where originally we are interpreting the world and that Where Art comes from music is purely so dove abstract or artistic period however, after I listen to the pieces, I realise that they really sound exactly like the instruments that are making the physical noises period for example it really sounds exactly like a physical piano period so I don't know about out a deep truth karma but it does seem that there is a physical sense that the music are represents which it can successfully mimic using this essentially image generating capability period one thing about all of these amazing AI development, is that I still make some long comments by dictating to Google. When it first got to the point that it was able to catch almost everything then was saying I was absolutely blown away period however, it's really not that good at taking dictation karma and I have to go back and replace each and every individual, and period with with the corresponding punctuation mark period seeing such an amazing developments happening month after month year after year ear makes me feel like we are really approaching what some people have called the singularity period when I read about out net positive fusion being announced my first Instinct was to think oh of course it's now that that chat GPT is available of course announcing a major fusion breakthrough would happen within in days to weeks it just makes perfect sense DJ eyes can solve problems that have have confounded scientists for decades period to see just how far we still have to go take a look at how this comment red before I manually corrected it to what I had actually set
- asdf333 4y agois classical music harder? noticed you didn't have any classical music tracks. i wonder if it is because it is more structured?
- TOMDM 4y agoThe audio sounds a bit lossy, would it be possible to create high quality spectograms from music, downsample them, and use that as training data for a spectogram upscaler? It might be the last step this AI needs to bring some extra clarity to the output.
- haykmartiros 4y ago/u/threevox on reddit made a colab for playing with the checkpoint: https://colab.research.google.com/drive/1FhH3HlN8Ps_Pr9OR6Qcfbfz7utDvICl0?usp=sharing https://colab.research.google.com/drive/1FhH3HlN8Ps_Pr9OR6Qc...
- nico 4y agoAmazing work. Can this be applied to voice? Example prompt: “deep radio host voice saying ‘hello there’” Kind of like a more expressive TTS?
- seth_ 4y agoAuthor here: It can certainly be applied to voice, but the model would need deeper training to speak intelligibly. If you want to hear more singing, you can try a prompt like "female voice", and increase the denoising parameter in the settings of the app. That said, our GPUs are still getting slammed today so you might face a delay in getting responses. Working on it!
- sergiotapia 4y agoReach out to the Beatstars CEO. He was looking for an AI play for his music producers marketplace. Probably solid B2B lead there.
- alsodumb 4y agoHayk! How smart are you! I loved your work on SymForce and Skydio - totally wasn't expecting you to be co-author on this! On a serious note, I'd really love some advice from you on time management and how you get so much done? I love Skydio and the problems you are solving, especially on the autonomy front, are HARD. You are the VP of Autonomy there and yet also managed to get this done! You are clearly doing something right. Teach us, senpai!
- poslathian 4y agoSuper! Makes sense since Skydio is also amazing. How much data is used for fine tuning? Since spectrograms are (surely?) very out of distribution for the pre training dataset, how much does value does the pre training really bring?
- haykmartiros 4y agoTo be honest, we're not sure how much value image pre training brings. We have not tried to train from scratch, but it would be interesting. One thing that's very important though is the language pre-training. The model is able to do some amazing stuff with terms that do not appear in our data set at all. It does this by associating with related words that do appear in the dataset.
- theGnuMe 4y agoYou can embed images in spectrograms.. might sound weird though
- picozeta 4y agoWhat I would really like to know - what happens if one trains that model from scratch (or is that not possible and training requirements are different? Sry for my ignorance, I never fine-tuned any diffusion model before)? In my experience (CNN based imagery segmentation) proven architectures (e.g. U-Net) performed similar with or without fine-tuning existing models (that have been mostly trained on imagenet, citiscapes, etc.) IF the domain was rather different. At least in the field of imagery segmentation there is not much of a point in fine-tuning an off-the-shelf model on let's say medical imagery. So maybe it's the same for the stable diffusion model. I don't see how some knowledge about the relationship between the prompt and given imagery describing that prompt should help this model map the prompt to a spectrogram of the given prompt.
- _leonie 4y agoHi, I really admire the skill you put at work on this project. At the same time, I think everyone is overlooking how crucial and problematic the training factor is. Why was stable diffusion able to generate spectrograms? Because it was fed some. Presumably, those original spectrograms were scraped with little concern over creators' permissions, just like it has been for artists' work in order to produce art-looking image generation. Please, research what has been happening in the art community lately. https://www.youtube.com/watch?v=Nn_w3MnCyDY https://www.youtube.com/watch?v=Nn_w3MnCyDY A protest on ArtStation has been shown to influence Midjourney's results, proving that huge amounts of proprietary work are constantly scraped without the creators' permission. AIs like these work so well just because they steal and remix real artists' work in the first place. There are going to be legal wars about this. Stable Diffusion doesn't have an official music generation Ai precisely because it couldn't train it with the same approach without being sued by music labels right away, while isolated artists don't have the same power. So, back to my question: have you wondered whose work is Stable Diffusion remixing here? Your endeavour is great technically, but as we progress into the future we have to be more aware of the ethical implications that come with different forms of progress. You could try to base your project on a collection of free-to-use spectograms, and see how it performs. If you do, I think it could actually be very interesting and useful to discuss the results here on Hacker News. Cheers!
- abledon 4y agoThis is groundbreaking! All other attempts at AI generated music have IMO, fallen flat... These results are actually listenable, and enjoyable! This is almost frightening how powerful this can be
- jablongo 4y agoHello - this is awesome work. Like other commenters, I think the idea that if you are able to transfer a concept into a visual domain (in this case via fft) it becomes viable to model with diffusion is super exciting but maybe an oversimplification. With that in mind, do you think this type of approach might work with panels of time series data?
- joelrunyon 4y agoThe site isn't working for me? Anything I have to fix on my side to make it work?
- plank 4y agoCrashes repeatedly on iOS in Firefox (my usual browser), is OK on Safari though, so probably not a webkit thing.
- tartakovsky 4y agoHi Hayk, I see that the inference code and the final model are open source. I am not expecting it, but is the training code and the dataset you used for fine-tuning, and process to generate the dataset open source?
- superkuh 4y agoI've compiled/run a dozen different image to sound programs and none of them produce an acceptable sound. This bit of your code alone would be a great application by itself. It'd be really cool if you could implement an MS paint style spectrum painting or image upload into the web app for more "manual" sound generation.
- juiiiced 4y agoDid you have a data set for training the relationship between words and the resulting sound?
- hanselot 4y agoThis is super awesome. Have you already explored doing the same with voice cloning?
- blaaaaa99a 4y agofunny that Hayk is an early skydio guy! 2 amazing AI projects. Huge respect :)
- ralfd 4y agoHow many songs did you use for the training data?
- rexreed 4y agoWhat sort of setup do you need to be able to fine tune Stable Diffusion models? Are there good tutorials out there for fine tuning with cloud or non-cloud GPUs?
- rexreed 4y ago"fine-tuned on images of spectrograms paired with text" How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.
- dqpb 4y agoObviously this needs a little more polish, but I've wanted this for so long I'm willing to pay for it now if it helps push the tech forward. Can I give you money?
- jweissman 4y agoThis is amazing! This is a fantastic concept generator. The verisimilitude with specific composers and techniques is more than a little uncanny. A few thoughts after exploring today… - My strongest suggestion is finding some strategy for smoothing over the sometimes harsh-sounding edge of the sample window - Perhaps it could be filling in/passing over segments of what is sounded to user as a larger loop? Both giving it a larger window to articulate things but maybe also showcasing the interpolation more clearly… - Tone control may seem challenging but I do wonder if you couldn’t “tune” the output of the model as a whole somehow (given the spectrogram format it could be a translation/scale knob potentially?)
- rhelsing 4y agoAmazing work! Do you plan on open-sourcing the code to train the model?
- d4rkp4ttern 4y agoSuper clever idea of course. But leaving aside how it was produced, I’ll be one of those who is underwhelmed by the musicality of this. I am judging this in terms of classical music. I repeatedly tried to get it to just play pure piano music without any other add-ons (cymbals etc). It kept mixing the piano with other stuff. Also the key question is - would something like this ever produce something as hauntingly beautiful and unique as classical music pieces?