37 ms·
Riffusion – Stable Diffusion fine-tuned to generate music
- corysama 4y agoCan anyone confirm/deny my theory that AI audio generation has been lagging behind progress in image generation because it’s way easier to get a billion labeled images than a billion labeled audio clips?
- madelyn 4y agoSound is a lot higher fidelity, it's harder to make the information available to a computer without serious downsampling or simplification. Consider sounds over 12khz. On a spectrogram during a chorus or drop that area is lit up, with so many things changing from millisecond to millisecond. A lot of AI samples really struggle at high frequencies, or even forgo them entirely. Midi based approaches have been really great though, and an approach like in the OP is fascinating (and impressive).
- antognini 4y agoAs someone who works in the audio AI space I think an underappreciated reason for audio lagging behind is that it's a lot harder to put impressive audio clips in an academic paper than impressive images.
- bufferoverflow 4y agoA network trained on spectrograms only should do much better.
- toasternz 4y agohttps://soundcloud.com/toastednz/stablediffusiontoddedwards?si=e7018bdda1014b8084d47d0ade6bf1ee&utm_source=clipboard&utm_medium=text&utm_campaign=social_sharing https://soundcloud.com/toastednz/stablediffusiontoddedwards?... 40 sec clip of uk garage/todd edwards style track made with riffusion -> serato studio with todd beats added
- gedy 4y agoThis is so good that I wondered if it's fake. Really impressive results from generated spectrographs! Also really interesting that it's not exactly trained on the audio files themselves - wonder if the usual copyright-based objections wild even apply here.
- deleted 4y ago[deleted]
- rbn3 4y agoregarding those usual objections, i'd argue that a spectrograph representation of a given piece of audio is just a different (lossy) encoding of the same content/information, so any hypothetical objections would still apply here.
- Applejinx 4y agoYou would be absolutely correct. the lossiness is in the resolution of the image (512x512 is pretty terrible) but given enough image resolution it's just an FFT transform, and the only reason that stuff falls short is because people don't give it, in turn, enough resolution. If you did wild overkill of the resolution of an FFT transform you could do anything you wanted with no loss of tone quality. If you turned that to visual images and did diffusion with it you could do AI diffusion at convincing audio quality. In theory the tone quality is not an objection here. When it sounds bad it's because it's 512x512, because the FFT resolution isn't up to the task, etc. People cling to very inadequate audio standards for digital processing, but you don't have to.
- kgwgk 4y agoWhy not? Music copyright was not even about audio recordings originally.
- talhof8 4y agoReally cool. Can't get this to work on the homepage though. Might be a traffic thing? Edit: Works now. A bit laggy but it works. Brilliant!
- deleted 4y ago[deleted]
- LoveMortuus 4y agoI also don't hear anything, even when my prompt was selected...
- MichaelZuo 4y agoMe neither, perhaps the web app is a bit buggy?
- scoopertrooper 4y agoI'm getting this back when I try to hear cats sing me a rock opera: {"data":{"success":true,"worklet_output":{"error":"Model version 5qekv1q is not healthy"},"latency_ms":530}}
- Pepe1vo 4y agoSame here, servers are overloaded probably. Shame, I was really looking forward to a Wu Tang Clan and Jamiroquai collab
- benplumley 4y agoSame earlier, but I can now get it to work very intermittently, with the error "Uh oh! Servers are behind, scaling up..."
- jansan 4y agoVery impressive. I am quite confident that next years number one Christmas hit will start like "church bells to electronic beats".
- quakeguy 4y agoXmd5a is already a real track. https://www.youtube.com/watch?v=crcqADcAusg https://www.youtube.com/watch?v=crcqADcAusg
- usrusr 4y agoRearrange that trip through latent space a little, jumping back and forth through different stages of the interpolation in a pattern resembling those customary chorus/verse things and you've got a hit. And you could reuse the exact same rearrange recipe for just about any interpolation between prompt pairs. Plenty of times this has been called before, and it's certainly possible that this wont be the last time this is called, but allow me to declare this the end of the bedroom producer (stage performers remain unaffected). And no two elevators will ever sound the same again, on any day.
- vikp 4y agoProducing images of spectrograms is a genius idea. Great implementation! A couple of ideas that come to mind: - I wonder if you could separate the audio tracks of each instrument, generate separately, and then combine them. This could give more control over the generation. Alignment might be tough, though. - If you could at least separate vocals and instrumentals, you could train a separate model for vocals (LLM for text, then text to speech, maybe). The current implementation doesn't seem to handle vocals as well as TTS models.
- btbuildem 4y agoI think you'd have to start with separate spectrograms per instrument, then blend the complete track in "post" at the end.
- rbn3 4y agogreat stuff, while it comes with the usual smeary iFFT artifacts that AI-generated sound tends to have the results are surprisingly good. i especially love the nonsense vocals it generates in the last example, which remind me of what singing along to foreign songs felt like in my childhood.
- dreilide 4y agoimpressive stuff. reminds me of when ppl started using image classifier networks on spectrograms in order to classify audio. i would not have thought to apply a similar concept for generative models, but it seems obvious in hindsight.
- superb-owl 4y agoThe interpolation from keyboard typing to jazz is incredible. This is what AI art should be.
- ZiiS 4y agoFor the 30 anniversary? https://warp.net/gb/artificial-intelligence https://warp.net/gb/artificial-intelligence
- hoherd 4y agoNice reference! I had never seen that site before, but those albums had a significant impact on my musical journey. There was a purple victorian house in Colorado Springs where the living room was converted into a record and cd store called Life By Design. I picked up these albums and a ton of other obscure music there. I was so happy to not have to drive all the way up to Wax Trax in Denver to be able to discover new artists.
- vintermann 4y agoFun! I tried something similar with DCGAN when it first came out, but that didn't exactly make nice noises. The conversion to and from Mel spectrograms was lossy (to put it mildly), and DCGAN, while impressive in its day, is nothing like the stuff we have today. Interesting that it gets so good results with just fine tuning the regular SD model. I assume most of the images it's trained on are useless for learning how to generate Mel spectrograms from text, so a model trained from scratch could potentially do even better. There's still the issue of reconstructing sound from the spectrograms. I bet it's responsible for the somewhat tinny sound we get from this otherwise very cool demo.
- knicholes 4y agoDoes anyone have any good guides/tutorials for how to fine-tune Stable Diffusion? I'm not talking about textual inversion or dreambooth.
- pea 4y agoThis is amazing! Would it be possible to use it to modify this interpolate between two existing songs (i.e. generate spectrograms from audio and transition between them)?
- minaguib 4y agoAbsolutely incredible - from idea to implementation to output.
- valdiorn 4y agoThis really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.
- 323 4y agoYou can also add another neural-network to "smooth" the spectrogram, increase the resolution and remove artefacts, just like they do for image generation.
- bckr 4y agoPretty sure that's how RAVE works
- hyperbovine 4y agoWasn't this Fraunhofer's big insight that led to the development of MP3? Human perception actually is pretty forgiving of perturbations in the Fourier domain.
- ComplexSystems 4y agoIn very limited situations. You can move a frequency around (or drop it entirely) if it's being masked by a nearby loud frequency. Otherwise, you would be amazed at the sensitivity of pitch perception.
- dehrmann 4y agoThe easy example of this is playing a slightly out of tune guitar, or a mandolin where the strings in the course aren't matched in pitch perfectly. You can hear it, and it's just a few cents off.
- w-m 4y ago
- bheadmaster 4y agoThis is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.
- josalhor 4y agoMakes me wonder if we will see a generalization of this idea. Just like in a CPU 90%+ of want you want to do can be modeled with very few instructions (mov, add, jmp..) we could see a set of very refined models (Stable difussion, GPT, etc) and all of their abstractions on top (ChatGPT, Rifussion, etc).
- DarmokJalad1701 4y agoMaybe next up is a model that generates Piet code https://www.dangermouse.net/esoteric/piet.html https://www.dangermouse.net/esoteric/piet.html
- chii 4y agoand you ask stable diffusion to generate piet code for a slightly better version of stable diffusion (or chatGPT) ...which then you can further use to generate a better version, and so on. Singularity here we come!
- danuker 4y agoIndeed, I think this would be a cost-effective way to go forward.
- amelius 4y agoPerhaps GPT could run on top of Stable-diffusion, generating output in the form of written text (glyphs).
- Tenoke 4y agoFor what is worth, people were trying the same thing with GANs (I also played with doing it with stylegan a bit) but the results weren't as good. The amazing thing is that the current diffusion models are so good that the spectograms are actually reasonable enough despite the small room for error.
- logn 4y agoCongratulations this is an amazing application of technology and truly innovative. This could be leveraged by a wide range of applications that I hope you'll capitalize on.
- m3kw9 4y agoThey’ve got a looooong way to go man
- ihatepython 4y agoI agree but it's better than listening to Ed Sheeran Edit: To be honest, I find something like 'Band In A Box' to be more impressive and actually useful, I don't understand how I would ever use this or listen to this. To me, it's further proof that Stable Diffusion really just doesn't work that well
- sexy_seedbox 4y agoDoesn't have enough training data for odd time signatures music... or John Cages' 4'33 :D
- nixpulvis 4y agoSounds a bit "clowny" to me, for lack of a better word.
- michpoch 4y agoEarlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?
- logn 4y agoThe raw outputs of these tools will be best consumed by experts. Until general AI, these are just better tools for the same workers.
- mensetmanusman 4y agoThey were killed off by the ability to record the data. Every city used to have their own music stars :)
- 323 4y agoPoliticians, bureaucracy. GPT-3, what policy should we apply to increase tax revenue by 5% given these constraints? GPT-3, please tell me some populist thing to say to win the next election, or how should I deflect these corruption charges.
- kmeisthax 4y ago"We should place a tax on all copyright lawyers and use it to fund GPU manufacturing and AI development. At your next stump speech, mention how the entertainment industry is stealing jobs from construction workers. Your corruption charges won't matter because voters only care about corruption when it's not in their favor."
- antipotoad 4y agoIsn’t this the plot of Deus Ex?
- Applejinx 4y agoMusicians were made to get a day job long before you were born ;)
- nathias 4y agoVery cool! I was wondering why there wasnt any music diffusion apps out there, it seems more useful because music has stricter copyright and all content creators need some background music ...
- lachlan_gray 4y agoWow, diffusion could be a game changer for audio restoration.
- kingcai 4y agoAbsolutely brilliant!
- adzm 4y agoReally fascinating. I'd be interested to know more about how it was trained, with what data exactly.
- deleted 4y ago[deleted]
- quux 4y agoThe vocals in these tracks are so interesting. They sound like vocals, with the right tone, phonemes. and structure for the different styles and languages but no meaning. Reminds me of the soundtrack to Nier Automata which did a similar thing: https://youtu.be/8jpJM6nc6fE https://youtu.be/8jpJM6nc6fE
- slenocchio 4y agoDo you guys think AI creative tools will completely subsume the possibility space of human made music? Or does it open up a new dimension of possibilities orthogonally to it? Hard for me to imagine how AI would be able to create something as unique and human as D'Angelo's Voodoo (esp. before he existed) but maybe it could (eventually). If I understand these AI algorithms at a high level, they're essentially finding patterns in things that already exist and replicate it w some variation quite well. But a good song is perfect/precise in each moment in time. Maybe we'll only be ever be able to get asymptotically closer but never _quite_ there to something as perfectly crafted a human could make? Maybe there will always be a frontier space only humans can explore?
- ElFitz 4y ago> Hard for me to imagine how AI would be able to create something as unique and human as D'Angelo's Voodoo (esp. before he existed) There’s always that immortal randomly typing monkey with a typewriter thing [1]. And, in our case, it seems to be better than random. So, yes, perhaps. But perhaps we could instead build and create things that are yet unimaginable upon it. We’ll see. [1]: https://en.wikipedia.org/wiki/Infinite_monkey_theorem https://en.wikipedia.org/wiki/Infinite_monkey_theorem
- rhelsing 4y agoI dont think its possible to make a compelling song in quite the same way. At Neptunely, we are pursuing a route that keeps the human in the picture. More of a collaboration. https://neptunely.com https://neptunely.com
- Applejinx 4y agoSome of this is really cool! The 20 step interpolations are very special, because they're concepts that are distinct and novel. It absolutely sucks at cymbals, though. Everything sounds like realaudio :) composition's lacking, too. It's loop-y. Set this up to make AI dubtechno or trip-hop. It likes bass and indistinctness and hypnotic repetitiveness. Might also be good at weird atonal stuff, because it doesn't inherently have any notion of what a key or mode is? As a human musician and producer I'm super interested in the kinds of clarity and sonority we used to get out of classic albums (which the industry has kinda drifted away from for decades) so the way for this to take over for ME would involve a hell of a lot more resolution of the FFT imagery, especially in the highs, plus some way to also do another AI-ification of what different parts of the song exist (like a further layer but it controls abrupt switches of prompt) It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do indistinct. It's realaudio grade, it needs to be more like 128kbps mp3 grade.
- Metus 4y ago> composition's lacking, too. It's loop-y. Well no wonder, it has absolutely no concept of composition beyond a single 5s loop, if I understand correctly. > It absolutely sucks at cymbals, though. Everything sounds like realaudio :) > It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do indistinct. It's realaudio grade, it needs to be more like 128kbps mp3 grade. I haven't sat down yet to calculate it, but is the output of SD at 512*512px at 24bit enough to generate audio CD quality in theory?
- TheOtherHobbes 4y agoNo. And I suspect this will always have phase smearing, because it's not doing any kind of source separation or individual synthesis. It's effectively a form of frequency domain data compression, so it's always going to be lossy. It's more like a sophisticated timbral morph, done on a complete short loop instead of an individual line. It would sound better with a much higher data density. CD quality would be 220500 samples for each five second loop. Realtime FFTs with that resolution aren't practical on the current generation of hardware, but they could be done in non-realtime. But there will always be the issue of timbres being distorted because outside of a certain level of familiarity and expectation our brains start hearing gargly disconnected overtones instead of coherent sound objects. What this is not doing is extracting or understanding musical semantics and reassembling them in interesting ways. The harmonies in some of these clips are pretty weird and dissonant, and not what you'd get from a human writing accessible music. This matters because outside of TikTok music isn't about 5s loops, and longer structures aren't so amenable to this kind of approach. This won't be a problem for some applications, but it's a long way short of the musical equivalent of a MidJourney image. Generally we're a lot more tolerant of visual "bugs" than musical ones.
- nonima 4y agoThis is really cool but can someone tell me why we are automating art? Who asked for this? The future seems depressing when I look at all this AI generated art.
- moonchrome 4y agoBecause art is the low hanging fruit of "close enough" applications.
- slenocchio 4y agoI wonder if this is true for music. Our ears are much more discerning than our eyes when it comes to art it seems.
- moonchrome 4y agoI mean listening to samples on the link above I'd hardly call it music so I'd say you're right.
- schwartzworld 4y agoYou can't automate a live performance or an oil painting with AI in this way. This isn't going to replace musicians and artists. If anything, I think a preponderance of AI art would make people appreciate the real stuff more. As to why, music is fun to create, and this is just a tool.
- sampo 4y ago> You can't automate a live performance or an oil painting with AI in this way. You'd have to combine it with these guys https://www.youtube.com/watch?v=WqE9zIp0Muk https://www.youtube.com/watch?v=WqE9zIp0Muk
- TheRealPomax 4y agoEveryone asked for this, including artists. If you make a living off of making art, having the best tools to help you do that is a constant, and the tools are finally starting to get properly good. Will "the job" change because of the tools? Of course. Will the nature of what it means for something to be art change? Also of course. Art isn't some static, untouchable thing. It changes as humanity does.
- bane 4y agoI bet a cool riff on this would be to simply sample an ambient microphone in the workplace and use that the generate and slowly introduce matching background music that fits the current tenor of the environment. Done slowly and subtly enough I'd bet the listener may not even be entirely aware its happening. If we could measure certain kinds of productivity it might even be useful as a way to "extend" certain highly productive ambient environments a la "music for coding".
- hammock 4y ago>in the workplace Or at a house party, club or restaurant... as more people arrive or leave and the energy level rises or declines..or human rhythms speed up or slow down...so does the music...
- Def_Os 4y agoDJs are getting automated away too!
- chrisfrantz 4y agoReactive generative music would so cool
- xwdv 4y agoOr perhaps use it in a hospital to play music that matches the state of a patient’s health as they are passing away.
- h4n1 4y agoI would not want to go to the hospital for a mild ear infection and hear the AI start blasting death metal.
- seth_ 4y agoAuthors here: Fun to wake up to this surprise! We are rushing to add GPUs so you can all experience the app in real-time. Will update asap
- deleted 4y ago[deleted]
- AMICABoard 4y agoAwesome, there is another project out there that does it with CPU https://github.com/marcoppasini/musika https://github.com/marcoppasini/musika maybe mix the both, ie take initial output of musika, convert to spectrogram and feed it to riffusion to get more variation...
- SamPatt 4y agoFascinating stuff. One of the samples had vocals. Could the approach be used to create solely vocals? Could it be used for speech? If so, could the speech be directed or would it be random?
- rexreed 4y ago"fine-tuned on images of spectrograms paired with text" How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.
- ubj 4y agoThis happened earlier than I expected, and using a much different technique than I expected. Bracing myself for when major record labels enter the copyright brawl that diffusion technology is sparking.
- flaviuspopan 4y agoI'm floored, the typing to jazz demo is WILD! Please keep pushing this space, you've got something real special here.
- Pepe1vo 4y agoI find it really cool that the "uncanny valley" that's audible on nearly every sample is exactly as I would imagine that the visual artifacts would sound that crop up in most generated art. Not really surprising I guess, but still cool that there's such a direct correlation between completely different mediums!
- epigramx 4y agoyeah, it's pretty unsurprising, that they're both uncanny valley like a messy circus.
- TuringNYC 4y agoI read the article: "If you have a GPU powerful enough to generate stable diffusion results in under five seconds, you can run the experience locally using our test flask server." Curious what sort of GPU the author was using or what some of the min requirements might be?
- genewitch 4y agoRTX 3070 can generate SD results in under 5 seconds, depending. Euler A 20 samples, 512x512. it can almost do 4 images in 5 seconds with those settings. It's possible a 3060 might work, depending. in my experience the 3060 is about 50% slower than the 3070, but that may be a bad 3060 in our test rig. but a 3060 gets pretty close to 5 seconds for an image, so try it, if you have one. just tested prompt "a test pattern for television" on both cards and 3070 took 1.87s and the 3060 took 2.93s. Similar results for the prompt "an intricate cityscape, like new york" edit: i should note we're using SD 1.4, not 1.5, although i think that just has to do with the checkpoint of the model, not the algorithm, but i could be wrong. Also the model is over 14GB, so perhaps the 3070 can't do it after all. i'll test it later as soon as the local admin wakes up and downloads it onto our machine.
- seth_ 4y agoAuthor here: fwiw we are running the app on a10g GPUs, which generally can turn around a 512x512 in 3.5s with 50 inference steps. This time includes converting the image into audio which should be done on the GPU as well for real-time purposes. We did some optimization such as a traced unet, fp16 and removing autocast. There are lots of ways it could be sped up further I'm sure!
- haykmartiros 4y agoOther author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the incredible results of stable diffusion, we were curious if we could fine tune the model to output spectrograms and then convert to audio clips. The answer to that was a resounding yes, and we became addicted to generating music from text prompts. There are existing works for generating audio or MIDI from text, but none as simple or general as fine tuning the image-based model. Taking it a step further, we made an interactive experience for generating looping audio from text prompts in real time. To do this we built a web app where you type in prompts like a jukebox, and audio clips are generated on the fly. To make the audio loop and transition smoothly, we implemented a pipeline that does img2img conditioning combined with latent space interpolation.
- jtode 4y agoAs one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.
- spoiler 4y agoAs a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a familiar word in a foreign language; the last one can happen more often though). So, I think these tools will be used to do background work (ie for visuals maybe help with background tasks in CGI or faraway textures in games). I know less about audio, but I assume it could maybe help a DJ create a transition between two segments they want to combine, as opposed make the whole composition for them, but idk if that example makes sense. Now, onto a more human point: I think that people often listen to music because it means something to them. Similar for people who appreciate visual art. I also love interactive and light art, and I love talking to other artists at light festivals who make them because of the stories and journeys behind their art too. Humans and art are a package deal, IMO. Edit: typos and to add: Also, I think prompt authorship is an art unto itself. I'm amazed what people can craft with it, but I'm more impressed by the craft itself than the outputs. Don't get me wrong, the outputs are darn cool, but not if you look closer. And it's impossible to look beneath the surface altogether, as there is nothing in the output but the pixels.
- joenot443 4y agoIncredible stuff, Seth & Hayk. I've been thinking nonstop about new things to build using Stable Diffusion and this is the first one that really took my breath away.
- xtracto 4y agoThis looks great and the idea is amazing. I tried with the prompt: "speed metal" and "speed metal with guitar riffs" and got some smooth rock-balad type music. I guess there was no heavy metal in the learning samples haha. Great work!
- the_third_wave 4y agoGregorian death metal folk also seems to have lacked seed tunes but the thing is just in its infancy so soon we'll be banging our tonsured heads to the folky beats of ... ...OK, need to create a band name generator to work in tandem with this thing. Let's see what one of its brethren in ML makes of it... - "Echoes of the Past": This name plays on the idea of Gregorian chanting, which is often associated with the distant past, and combines it with the intense and aggressive sound of death metal. - "The Order of the Black Chant": This name incorporates elements of both the religious connotations of Gregorian chanting and the dark, heavy sound of death metal, creating a sense of mystery and danger. - "Foretold in Blood": This name evokes both the ancient, mystical nature of Gregorian chanting and the violent themes of death metal, creating a sense of ancient prophecy coming to pass. - "Crypt of the Silent Choir": This name brings together the eerie, otherworldly sound of Gregorian chanting with the underground, underground feel of death metal, creating a sense of hidden secrets and forbidden knowledge. "The Order of the Black Chant" it shall be.
- cwillu 4y agoSomewhat unrelated, but given the descriptions, perhaps you should take a listen to the Darktide 40k soundtrack if you haven't: https://www.youtube.com/watch?v=D4hEOMSjzdo https://www.youtube.com/watch?v=D4hEOMSjzdo
- the_third_wave 4y agoYes, that sounds like something which goes in the right direction. Now just drop the modern stuff - the synth bass beat doesn't really fit the description - and have some more Gregorian growl and Lucifer's your uncle.
- 4y ago
- sampo 4y agoGPT-3 has 175 billion parameters (says Wikipedia). What is the size of the neural network used in this riffusion project?
- sebzim4500 4y agoStable Diffusion 'only' has ~1B parameters IIRC.
- senko 4y ago@haykmartiros, @seth_, thank you for open sourcing this! Played a bit with the very impressive demos, now waiting in queue for my very own riff to get generate. Great as this is, I'm imagining what it could do for song crossfades (actual mixing instead of plain crossfade even with beat matching).
- wwarner 4y agoBOOM! Yes!
- CrypticShift 4y agoThings similar to the “interpolation” part (not the generative part) are already used extensively especially for game and movie sound design. Kyma [1] is the absolute leader (it requires expensive hardware though). IMO later iterations on this approach may lead to similar or better results. FYI, other apps that use more classic but still complex Spectral/Granular algos : https://www.thecargocult.nz/products/envy https://www.thecargocult.nz/products/envy https://transformizer.com/products/ https://transformizer.com/products/ https://www.zynaptiq.com/morph/ https://www.zynaptiq.com/morph/ [1] https://kyma.symbolicsound.com/ https://kyma.symbolicsound.com/
- serverholic 4y agoI’m curious about the limitations of using spectrograms and transient-heavy sounds like drums. It seems like you’d need very high resolution spectrograms to get a consistently snappy drum sound.
- genewitch 4y ago8GB is enough to do 1080p resolution. the UI i use for SD maxes out at 2048x2048. however, it takes a lot longer than 512x512 to generate: 1m40s versus 1.97s. I'm guessing if one had access to one of those nvidia backplane rackmount devices one could generate 8k or larger resolution images.
- astrange 4y agoSD can’t generate coherent images if you increase the output size. They’re basically always unusable unless you don’t need any global architecture to them.
- genewitch 4y agoI'm not sure what you mean. with the 1.4 checkpoint i can make a scene "a sky full of blimps" and then img2img that into a space battle where all the blimps in the sky are either fires or spacecraft afterburners. at 2k pixels, and then use the built in "GAN" stuff to bring it to 4k, 5k, or 8k. It isn't perfect and it takes a lot of fiddling and sysadmin stuff, but it is way better than nextchar = CHR$(RAND64); color = (rand65535); of yore.
- deleted 4y ago[deleted]
- xcambar 4y agoI will try it but at least for the name it deserves praise.
- 2devnull 4y agoWas just watching an interview of Billy Corgan (smashing pumpkins) on Rick Beato’s YouTube[1] last night where billy was lamenting the inevitable future where the “psychopaths” in the music biz will use ai and auto tune to churn out three chord non-music mumble rap for the youth of tomorrow, or something to that effect. It was funny because it’s the sad truth. It’s already here but new tech will allow them to cut costs even more, and increase their margins. No need for musicians. Really cool on one hand, in the same way fentanyl is cool — or the cotton gin, but a bit depressing on the other, if you care about musicians. I and a few others will always pay to go the symphony, so good players will find a way get paid, but this is what kids will listen to, because of the profit margin alone. [1] https://m.youtube.com/watch?v=nAfkxHcqWKI https://m.youtube.com/watch?v=nAfkxHcqWKI
- bogwog 4y agoDamn this is insane. I wonder what other things can be encoded as images and generated with SD?
- d7y 4y ago
- andy_ppp 4y agoI was thinking about this - what if someone trained a stable diffusion type model on all of the worlds commercial music? This model would probably produce quite amazing music given enough prompting and I'm wondering if the music industry would be allowed to claim copyright on works created with such a model. Would it be illegal or is this just like a musician picking up ideas from hearing the world of music? Is it really right to make learning a crime, even if machines are doing it? I'm conflicted after finding out that for sync licensing the music industry want a percentage of revenue based on your subscription fees, sometimes as high as 15%-20%! I'm surprised such a huge fee isn't considered some kind of protection racket.
- throw78311 4y agoThis question has been explored before, see Kologorov Music: https://www.youtube.com/watch?v=Qg3XOfioapI https://www.youtube.com/watch?v=Qg3XOfioapI
- londons_explore 4y agoI think there has to be a better way to make long songs... For example, you could take half the previous spectrogram, shift it to the left, and then use the inpainting algorithm to make the next bit... Do that repeatedly, while smoothly adjusting the prompt, and I think you'd get pretty good results. And you could improve on this even more by having a non-linear time scale in the spectrograms. Have 75% of the image be linear, but the remaining 25% represent an exponentially downsampled version of history. That way, the model has access to what was happening seconds, minutes, and hours ago (although less detail for longer time periods ago).
- someguyorother 4y agoPerhaps you could do a hierarchical approach somehow, first generating a "zoomed out" structure, then copying parts of it into an otherwise unspecified picture to fill in the details. But perhaps plain stable diffusion wouldn't work - you might need different neural networks trained on each "zoom level" because the structure would vary: music generally isn't like fractals and doesn't have exact self-similarity.
- mdonahoe 4y agoYou seem smart. How do I follow you?
- winReInstall 4y agoCant wait to see this in karaoke, you just sing lyrics and it jams along with music.
- zone411 4y agoInteresting. I experimented a bit with the approach of using diffusion on whole audio files, but I ultimately discarded it in favor of generating various elements of music separately. I'm happy with the results of my project of composing melodies (https://www.youtube.com/playlist?list=PLoCzMRqh5SkFPG0-RIAR8jYRaICWubUdx https://www.youtube.com/playlist?list=PLoCzMRqh5SkFPG0-RIAR8...) and I still think this is the way to go and but that was before Stable Diffusion came out. These are interesting results though, maybe it can lead to something more.
- bluebit 4y agoAnd we broke it.
- londons_explore 4y agoI propose that while you are GPU limited, you make these changes: * Don't do the alpha fade to the next prompt - just jump straight to alpha=1.0. * Pause the playback if the server hasn't responded in time, rather than looping.
- TechTechTech 4y agoI got an actual `HTTP 402: PAYMENT_REQUIRED` response (never seen one of those in the wild, according to Mozilla it is experimental). Someone's credit card stopped scaling?
- haykmartiros 4y agoLOL. Yes we had to upgrade our Vercel tier: https://twitter.com/sethforsgren/status/1603425188401467392 https://twitter.com/sethforsgren/status/1603425188401467392
- newswasboring 4y agoIf it can do music, can we train better models for different kinds of music? Or different models for different instruments makes more sense? For different instruments we can get better resolution by making the spectrogram represent different frequency ranges. This is terribly exciting, what a time to be alive.
- fritzschopen 4y agoit seems that SD does cover everything in terms of generative ai. Speaking of music, very interesting paper and demo. Just wondering in terms of license and commercialization, what kind of mess are we expecting here?
- Abecid 4y agoThis is one of the most ingenious thing I've seen in my life
- r3trohack3r 4y agoI can’t help but see parallels to synesthesia. It’s amazing how capable these models are at encoding arbitrary domain knowledge as long as you can represent it visually w/ reasonable noise margins.
- woeirua 4y agoVery cool, but the music still has a very "rough", almost abrasive tinge to it. My hunch, is that it has to do with the phase estimates being off. Who's going to be first to take this approach and use it to generate human speech instead?
- rslice 4y agodeleted
- dangond 4y agoI think you'll find plenty of people who find that DAWs and music theory help them better find self-expression and celebrate life through their music. Any tool or framework that opens up new modes of achieving that self-expression should be celebrated, not shunned because it isn't as "pure" as more time and labor intensive methods. Would you rather someone be forced to dedicate a significant amount of time to studying music and art creation just to be able to find that self-expression?
- deleted 4y ago[deleted]
- gitfan86 4y agoThis is what I've been talking about all year. It is such a relief to see it actually happen. In summary: The search for AGI is dead. Intelligence was here and more general than we realized this whole time. Humans are not special as far as intelligence goes. Just look how often people predict that an AI cannot X or Y or Z. And then when an AI does one of those things they say, "well it cannot A or B or C". What is next: This trend is going to accelerate as people realize that AI's power isn't in replacing human tasks with AI agents, but letting the AI operate in latent spaces and domains that we never even thought about trying.
- visarga 4y agoGenerated contents without filtering/validation are worthless. I predict some kind of testing, validation or ranking be developed to filter out generated contents. Each domain has its own rules - you need to implement validation for code and math, fact checks for text, contrasting the results from multiple solutions for problem solving, and aesthetic scoring for art. But validation is probably harder than learning to generate in the first place, probably a situation similar to closing the last percent in self driving.
- mastax 4y agoWow those examples are shockingly good. It's funny that the lyrics are garbled analogously to text in stable diffusion images. The audio quality is surprisingly good, but does sound like it's being played through an above-average quality phone line. I bet you could tack on an audio-upres model afterwards. Could train it by turning music into comparable-resolution spectrograms.
- tomrod 4y agoThis is huge. This show me that Stable Diffusion can create anything with the following conditions: 1. Can be represented as as static item on two dimensions (their weaving together notwithstanding, it is still piece-by-piece statically built) 2. Acceptable with a certain amount of lossiness on the encoding/decoding 3. Can be presented through a medium that at some point in creation is digitally encoded somewhere. This presents a lot of very interesting changes for the near term. ID.me and similar security approaches are basically dead. Chain of custody proof will become more and more important. Can stable diffusion work across more than two dimensions?
- marviel 4y agoI would argue that its high-fidelity representations of 3d space, imply that the model's weights are capable of pattern-matching in multiple dimensions, provided the input is embedded into 2d space appropriately.
- Pxtl 4y agoNow I'm wondering about feeding Stable Diffusion 2D landscape data with heightmaps and letting generate maps for RTS videogames. I mean, the only wrinkle there is an extra channel or two.
- emporas 4y agoAny image generator can do well in any two dimensional data, including SD, Dalle, Midjourney. One feature of SD not discussed much in my opinion, is the deterministic key it provides to the user. This is what enables the smooth transition in every second of music it generates, and the next second in time. Moving the cursor of latent space between in a minimal way, creating the next piece of information and change it ever so slightly, it definitely sounds good to the human ears.
- dwringer 4y agoBeing able to blend between prompts and attention weightings smoothly from a fixed seed is definitely a fantastic and underexplored avenue; it makes me recall "vector synthesis" common in wavetable synthesizers since the '80s as discussed here[0]. I feel we are just a couple of months from seeing people start using MIDI controllers to explore these kinds of spaces. Something could be hacked together today, but it will be interesting to see once the images can be generated in nearly realtime as the controls are adjusted. [0] https://www.soundonsound.com/techniques/synth-school-part-7 https://www.soundonsound.com/techniques/synth-school-part-7
- gardenhedge 4y agoimpressive. and this is a hobby project.. amazing
- EZ-Cheeze 4y ago"https://en.wikipedia.org/wiki/Spectrogram https://en.wikipedia.org/wiki/Spectrogram - can we already do sound via image? probably soon if not already" Me in the Stable Diffusion discord, 10/24/2022 The ppl saying this was a genius idea should go check out my other ideas
- hoschicz 4y agoWhat did use as training data?
- 451mov 4y agowhy not use an image of the waveform as input?
- naillo 4y agoIt's interesting if this can be used for longer tracks by inpainting the right half of the spectrogram.
- stevehiehn 4y agoReally great! I've been using diffusion as well to create sample libraries. My angle is to train models strictly on chord progression annotated data as opposed to the human descriptions so they can be integrated into a DAW plugin. Check it out: https://signalsandsorcery.org/ https://signalsandsorcery.org/
- wmwmwm 4y agoThis is amazing, and scary (as a musician) but also reliably kills firefox on iOS!
- spyder 4y agoAnother related audio diffusion model (but without text prompting) here: https://github.com/teticio/audio-diffusion https://github.com/teticio/audio-diffusion
- kanwisher 4y agooh wow this one works really well
- lucidrains 4y agopersonalized RL agents that finds aesthetic trajectories through the music latent space... soon, i hope :D
- haykmartiros 4y agoLove this idea. If I had more time I wanted to make a spaceship game where you are flying around the latent space, and model interrogation is used to provide labels to landmarks as you move around.
- Animats 4y ago"Uh oh! Servers are behind, scaling up..." - havent' been able to get past that yet. Anyone getting new output? This is already better than most techno. I can see DJs using this, typing away.
- NHQ 4y agoIn the end there was the word.
- evo_9 4y agoPretty nice, I was just talking to a friend about needing a music version of chatgpt, so thank you for this. Wondering if it would be possible to create a version of this that you can point at a person SoundsCloud and have it emulate their style / create more music in the style of the original artist. I have a couple albums worth of downtempo electronic music I would love to point something like this at and see what it comes up with.
- trekkie1024 4y agohttps://mubert.com/ https://mubert.com/ might be what you're looking for.
- evo_9 4y agoThank you, I will check it out!
- rhelsing 4y agoWe are working on something similar at Neptunely. https://neptunely.com https://neptunely.com
- birdyrooster 4y agoI know it sounds like I am going to be sarcastic, but I mean all of this in earnest and with good intention. Everything this generates is somehow worse than the thing it generated before it. Like the uncanny valley of audio had never been traversed in such high fidelity. Great work!
- deleted 4y ago[deleted]
- pmontra 4y agoA musician friend of mine told me that this is (I freely translate) a perversion, building in frequency and returning time. Don't shoot the messenger. Personally I like the results. I'm totally untrained and couldn't hear any of the issues many comments are pointing out. I guess that all of lounge/elevator music and probably most ad jingles will be automated soon, if automation cost less than human authors.
- adamsmith143 4y ago"Horse Carriage Driver says horseless carriages are abominations. More at 12!"
- tltimeline2 4y agoadamsmith143, commenter. Witness here the apathy of someone unaffected by the suffering of millions and uncommitted to the soul of humanity. DIAF.
- adamsmith143 4y agoSorry my sympathy for "starving" artists isn't sufficient for your liking. 95% of modern art is drivel that can and should be replaced by AI art. A recent walk through the MoMa in NY had me weeping for humanity. Filled with, in some cases literal, trash that any 5th grader could produce.
- dylan604 4y agoThe results of this are similar to my nitpicks of AI generated images (well, duh!). There's definitely something recognizable there, but somethings just not quite right about it. I'm quite impressed that there was enough training data within SD to know what a spectrograph looks like for the different sounds.
- Broge 4y agoI wonder if it's possible to fine-tune an image upscaling model on spectrograms, in order to clean up the sound?
- soperj 4y ago> https://www.riffusion.com/?&prompt=punk+rock+in+11/8 https://www.riffusion.com/?&prompt=punk+rock+in+11/8 Tried getting something in an odd timing, but still is 4/4.
- Aardwolf 4y agoHow comes that the stable diffuse model helps here? Does the fact that it knows what an astronaut on a horse looks like have effect on the audio? Would starting the training from an empty model work too?
- neontomo 4y agoI think they only mentioned the horse to illustrate to people that they are using the same tool as what was used to generate those types of images. It's painting a picture for the uninitiated audience. From what I understood, this model is trained on spectrograms instead of horses and the like, resulting in this product.
- aquanext 4y agoSomeone please train it on John Coltrane.
- lftl 4y agoI just wanted to say you guys did an amazing job packaging this up. I managed to get a local instance up and running against my local GPU in less than 10 minutes.
- mensetmanusman 4y agoThis works because songs are images in time. FFT analysis does not care.
- simsspoons 4y agothis is just great
- zoytek 4y agoIt's amazing. They've really got something revolutionary here.
- isoprophlex 4y agoI wonder how they got their train data..! The spectrogram trick is genius, but not much useful without high quality, diverse data to train on
- jsat 4y agoToday's music generation is putting my Pop Ballad Generator to shame: http://jsat.io/blog/2015/03/26/pop-ballad-generator/ http://jsat.io/blog/2015/03/26/pop-ballad-generator/
- bulbosaur123 4y agoAnyone interested in joining an unofficial Riffusion Discord, let's organize here: https://discord.gg/DevkvXMJaa https://discord.gg/DevkvXMJaa Would be nice to have a channel where people can share Riffs they come up with.
- PcChip 4y agothe problem is it sounds awful, like a 64kbps MP3 or worse Perhaps AI can be trained to create music in different ways than generating spectrograms and converting them to audio?
- w_for_wumbo 4y agoIt doesn't need to sound good at all for it to be useful. Like with the AI Art creation, it can be a starting point for artists to play around and rapidly try different concepts, and then interpret the concept using high quality tools to create something really quite remarkable. It's all about empowering artists to explore more possibilities.
- stevehiehn 4y agoExactly! UX/Pipelines/Integrations are the next logical step. It's my belief that samples will essentially be 'free' very soon. We will see DAW plugins/integrations that contextually offer samples to the composer. I'm confident in this because that's what I'm working on.
- magicalhippo 4y agoI haven't been paying close attention to the AI field but seeing someone painting with NVIDIA Canvas blew me away[1]. I can totally see how it can help prototyping or exploring ideas and visions. [1]: https://www.youtube.com/watch?v=uLlhyKxygrI https://www.youtube.com/watch?v=uLlhyKxygrI
- ricopags 4y agoThis is so completely wild. Love the novelty and inventiveness. Could anyone help me understand whether using SVG instead of bitmap image would be possible? I realize that probably wouldn't be taking advantage of the current diffusion part of Stable-Diffusion, but my intuition is maybe it would be less noisy or offer a cleaner/more compressible path to parsing transitions in the latent space. Great idea? Off base entirely? Would love some insight either way :D
- rmetzler 4y ago"Jamaican rap" - usually the genre (e.g. Sean Paul) is called Dancehall.
- raajg 4y agoIf such unreasonably good music can be created based on information encoded in an image, I'm wondering what there things we can do with this flow: 1) Write text to describe the problem 2) Generate an image Y that encodes that information 3) Parse that image Y to do X Example: Y = blueprint, X = Constructing a building with that blueprint
- orobinson 4y agoI’d been wondering (naively) if we’d reached the point where we can’t see any new kinds of music now that electronic synthesis allows us to make any possible sound. Changes in musical styles throughout history tend to have been brought about by people embracing new instruments or technology. This is the most exciting thing I’ve seen in ages as it shows we may be on the verge of the next wave of new technology in music that will allow all sorts of weird and wonderful new styles to emerge. I can’t wait to see what these tools can do in the hands of artists as they become more mainstream.
- EamonnMR 4y ago'make any possible sound' is less important than 'make x sound easily' by way of tools and accumulated knowledge. Also what's audiences are receptive to matters a lot - you could have made noise rock in the 40s but I can't imagine it would have sold a lot of records.
- GaggiX 4y agoYou can train/finetuned a Stable Diffusion model on an arbitrary aspect ratio/resolution and then the model starts creating coherent images, would be cool to try finetuning/training this model on entire songs by extending the time dimension (also the attention layer at the usual 64x64 resolution should be removed or it would eat too much memory)
- fernandohur 4y agoI found this awesome podcast that goes into several AI & music related topics https://open.spotify.com/show/2wwpj4AacVoL4hmxdsNLIo?si=IAaJ6A3yQ1GT5na4iSZnRA https://open.spotify.com/show/2wwpj4AacVoL4hmxdsNLIo?si=IAaJ... They even talk specifically about about applying stable diffusion and spectrograms.
- Raed667 4y agoSeems to be victim of its own success: - No new result available, looping previous clip - Uh oh! Servers are behind, scaling up I hope Vercel people can give you some free credits to scale it up.
- motoxpro 4y agoThis is just insane. Sooooo incredible. Don't really realize how far things have come until it hits a domain you're extremely familiar with. Spent 8-9 in music production and the transition stuff blew me away.
- needz 4y agoThis website crashes Firefox on iOS
- Moosdijk 4y agoWow this is awesome!
- phneutral26 4y agoRight now it still seems to lack the horsepower for this many users. Hope it gets in a better state soon, but I am bookmarking this right now!
- esotericsean 4y agoComing at this from a layman's perspective, would it be possible to generate a different sort of spectrogram that's designed for SD to iterate upon even more easily?
- leod 4y agoAwesome work. Would you be willing to share details about the fine-tuning procedure, such as the initialization, learning rate schedule, batch size, etc.? I'd love to learn more. Background: I've been playing around with generating image sequences from sliding windows of audio. The idea roughly works, but the model training gets stuck due to the difficulty of the task.
- owlbynight 4y agoIf copyright laws don't catch up, the sampling industry is cooked. Made this: https://soundcloud.com/obnmusic/ai-sampling-riffusion-waves-lapping-on-a-shore-to-nowhere https://soundcloud.com/obnmusic/ai-sampling-riffusion-waves-...
- navane 4y agoassuming the first sample was generated by OP method, how did you clean that sample up so nicely?
- owlbynight 4y agoYep, the first sample was generated via OP's method. The tl;dr answer is that I used a few plugins in my digital audio workstation to make it sound better. I made a video specifically in reply to you if you want to see exactly how I did it (3 mins): https://www.youtube.com/watch?v=69Q-cseNCI4 https://www.youtube.com/watch?v=69Q-cseNCI4
- navane 4y agoThank you so much!
- ElijahLynn 4y agoWow, I just learned so much about spectograms, had no idea that one could reverse one into audio waves!
- ElijahLynn 4y agoI was confused because I must not have read good that the working webapp is at https://www.riffusion.com/ https://www.riffusion.com/. Go to https://www.riffusion.com/ https://www.riffusion.com/ and press the play button to see it in action!
- XorNot 4y agoSo this is slightly bending my mind again. Somehow image generators were more comprehensible compared to getting coherent music out. This is incredible.
- fowlkes 4y agoMultiple folks have asked here and in other forums but I'm going to reiterate, what data set of paired music-captions was this trained on? It seems strange to put up a splashy demo and repo with model checkpoints but not explain where the model came from... is there something fishy going on?
- up2isomorphism 4y agoThese are horrible musics, but of course there is nothing to feel shame about it.
- farmin 4y agoThat church bell one is amazing. Very creative transition.
- MagicMoonlight 4y agoPlug this into a video game and you could have GTA 6 where the NPCs have full dialogue with the appearance of sentience, concerts where imaginary bands play their imaginary catalogue live to you and all kinds of other dynamically generated content.
- Slow_Hand 4y agoAs a musician, I'll start worrying once an AI can write at the level of sophistication of a Bill Withers song: https://www.youtube.com/watch?v=nUXgJkkgWCg https://www.youtube.com/watch?v=nUXgJkkgWCg Not simply SOUND like a Bill Withers song, but to have the same depth of meaning and feeling. At that point, even if we lose we win because we'll all be drowning in amazing music. Then we'll have a different class of problem to contend with.
- bawolff 4y agoI wonder if this would be applicable to video game music. Be able to make stuff that's less repetitive but also smoothly transitions to specific things with in-game events.
- MarcelOlsz 4y agoNow if only it could generate accurate sheet music, then you've got something. Incredible examples.
- vermilingua 4y agoFinally, post-avant jazzcore [1] and progressive dreamfunk [2] made real. [1] https://www.riffusion.com/?&prompt=Post-avant+jazzcore&seed=8730&denoising=0.75&seedImageId=og_beat https://www.riffusion.com/?&prompt=Post-avant+jazzcore&seed=... [2] https://www.riffusion.com/?&prompt=Progressive+dreamfunk&seed=3&denoising=0.75&seedImageId=og_beat https://www.riffusion.com/?&prompt=Progressive+dreamfunk&see...
- Scaevolus 4y agoDid you fine-tune the VAE?
- sliss 4y agoSuch a creative application of SD to spectrograms! ...now do stock charts
- billiam 4y agoIt may be clearer to those of you who are smarter than me, but I guess I've only recently begun to appreciate what these experiments show--that AI graphical art, literature, music and the like will not succeed in lowering the barriers to humans making things via machines but in training humans to respond to art that is generated by machines. Art will not be challenging but designed by the algorithm to get us to like it. Since such art can be generated for essentially no cost, it will follow a simple popularity model, and will soon suck like your Netflix feed.
- imiric 4y ago> Since such art can be generated for essentially no cost, it will follow a simple popularity model, and will soon suck like your Netflix feed. I'm not so sure. Considering how successful AI-driven social media feeds are, which already include substantial AI-generated content, why would a feed consisting entirely of such content be any less successful? The quality will only keep increasing. > Art will not be challenging but designed by the algorithm to get us to like it. I don't think these advancements are a threat to art created by humans, just as any art created by humans isn't a threat to other art. It's just... more art. Eventually, AI will be capable of being truly creative, instead of being trained on human art and producing permutations of it, which will also be wonderful. The role of humans will be to train these models to produce art we find enjoyable. Imagine if your AI media feed was an infinite stream of artworks personalized just for your taste. It will be TikTok on steroids. I can't say I'm thrilled by that prospect, because it will also be used for exploiting users, but the entertainment potential is huge.
- analog31 4y agoRight now, the AI is trained on mostly human generated art. It will be interesting to see what happens when the training set itself is mostly AI generated. The role of the "artist" in the future will not be to create art directly, but to influence art by manipulating the training set. For instance an artist who can flood the Internet with a trillion captioned images will be able to spawn an art movement, even if it short lived.
- fire 4y agounfortunately I put in "sonic the hedgehog" and the result was... ... nothing? like, there's no playback at all. Is that expected?
- primarch 4y agoCould this concept be inverted, to identify music from a text prompt, as in i want a particular vibe and it can go tell me what music fits that description. Ive always thought the ability to find music you like was very lacking, it should not be bounded by a genre instead its usually based on rythmic and melodic structures that appeal to you, regardless of what type of music it might be.
- xor99 4y agoAlso check the similar work on arxiv: Multi-instrument Music Synthesis with Spectrogram Diffusion: https://arxiv.org/abs/2206.05408 https://arxiv.org/abs/2206.05408
- johndhi 4y agoI'm a short fiction writer. Do you think I could get one of these new models to write a good story? I'd want to train it to include foreshadowing, suspense, relatable characters and perhaps a twist ending that cleverely references the beginning.
- burlesona 4y agoI’ve thought about this and tested it a lot, it doesn’t work. The reason is, the models don’t understand content or even context, they just recognize patterns and can generate similar patterns which we then interpret. Case in point, a photo of an Astronaut on a horse is not actually an astronaut on a horse, it’s a 2D pixel map of light that our eyes then interpret to mean an astronaut on a horse. It’s even easier to understand this listening to the generated singing. It sounds just like singing - but it’s not language at all and doesn’t mean anything, it’s just a very similar pattern. These AI models are great at making things that look like a pattern that to us means something, but they don’t generate semantically meaningful content itself. So when you do this with GPT3 or other models and long-form narrative it falls apart pretty fast as the model can’t keep straight things like characters and their internal personalities and motives, nor the overall plot arc. But! You can definitely feed in prompts and ask questions to get ideas and boilerplate - descriptions of people and places come out especially well - and then you can edit that and use it to accelerate your writing process.
- RugnirViking 4y agothey can only do about 150 words at a time, so you'd struggle to get anything longform out of it. You can keep asking it over and over but then its liable to forget previous information. You'd do better with a prompt that repeatedly reminds the AI the style its going for and some basic information about the characters, but it's story would likely be quite cliche in many ways. It's something i've experimented with trying to get it to make DND adventure books
- rhelsing 4y agoI think it is possible. Have you tried interacting with chatGPT on short story?
- LonelyWolfe 4y agoI wonder if subliminal messaging will somehow make a comeback once we have ai generated audio and video. Something like we type "Fun topic" and those controlling the servers will inject "and praise to our empire/government/tyrant" to the suggestion or something like that.
- Fricken 4y agoThis opens up ideas. One thing people have tried to do with stable diffusion is create animations. Of course, they all come out pretty janky and gross, you can't get the animation smooth. But what if what if a model was trained not on single images, but animated sequential frames, in sets, laid out on a single visual plane. So a panel might show a short sequence of a disney princess expressing a particular emotion as 16 individual frames collected as a single image. One might then be able to generate a clean animated sequence of a previously unimagined disney princess expressing any emotion the model has been trained on. Of course, with big enough models one could (if they can get it working) produced text prompted animations across a wide variety of subjects and styles.
- gavindean90 4y agoWe are back to sprite maps
- nathan_f77 4y agoThat's an interesting idea. I wonder if this would work with inpainting - erase the 16th cell and let the AI fill it in. Then upscale each frame. Has anyone experimented with this?
- zamadatix 4y agohttps://www.reddit.com/r/StableDiffusion/comments/yj1kbi/ive_trained_a_new_model_to_output_pixel_art/ https://www.reddit.com/r/StableDiffusion/comments/yj1kbi/ive...
- Fricken 4y agoWell Look at that. I'm totally not surprised, lol.
- tekakutli 4y agoa different approach: analog television
- xor99 4y agoNew era of library music.
- ngcc_hk 4y agoNot generate … steal.
- deleted 4y ago[deleted]
- nathan_f77 4y agoThis is awesome! It would be interesting to generate individual stems for each instrument, or even MIDI notes to be rendered by a DAW and VST plugins. It's unfortunate that most musicians don't release the source files for their songs so it's hard to get enough training data. There's a lot of MIDI files out there but they don't usually have information about effects, EQ, mastering, etc.
- rhelsing 4y agoLet's get in touch. This is precisely what we are working on at Neptunely. https://neptunely.com https://neptunely.com
- NoPicklez 4y agoThis, Chat-GPT and the AI image generation. We're now at a very interesting time where average joes get to start using incredible tools.
- fillipvt 4y agoNow the next step is to use Stable Diffusion to create chemical components via molecular graphs
- Doorstep2077 4y ago
- wereallterrrist 4y ago:/ this is just spam. How are more of your comments not flagged?
- Doorstep2077 4y agonice username
- r-k-jo 4y agoAmazing project! Here is a demo including an input for negative prompt. It's impressive how it works. You can try: prompt: relaxing jazz melody bass music negative_prompt: piano music https://huggingface.co/spaces/radames/spectrogram-to-music https://huggingface.co/spaces/radames/spectrogram-to-music
- epigramx 4y agoI wasn't expecting to see uncanny valley translated to music today.
- blaaaaa99a 4y agoI feel like the next step here is to get a GPT3 like model to parse everything ever written about every piece of music which is on the internet (and in every pdf on libgen and scihub) and link them to spectrograms of that music and then things are going to get wild I am so blessed to live in this era :)
- O__________O 4y ago>> Prompt - When providing prompts, get creative! Try your favorite artists, instruments like saxophone or violin, modifiers like arabic or jamaican, genres like jazz or rock, sounds like church bells or rain, or any combination. Many words that are not present in the training data still work because the text encoder can associate words with similar semantics. The closer a prompt is in spirit to the seed image And BPM, the better the results. For example, a prompt for a genre that is much faster BPM than the seed image will result in poor, generic audio. (1) Is there a corpse the keywords were collected from? (2) Is it possible to model the proximity of the image to keywords and sets of keywords?
- whiddershins 4y agoThis is a brilliant idea. Also, spectrographs will never generate plausible high quality audio. (I think) So I think the next move is to map the generate audio back over to synthesizer and samples via midi …
- realmod 4y agoWow, absolutely fascinating. AI will continue to revolutionize our current approaches.
- _spduchamp 4y agoI prefer to make my spectrograms by hand. https://youtu.be/HT0HH_fc4ZU https://youtu.be/HT0HH_fc4ZU
- throwaway743 4y agoSome how it made Wesley Willis sound even better.
- sagebird 4y agoIs there a different mapping of FFT information to a two dimensional image that would make harmonic relationships more obvious? IE, use a polar coordinate system where angle 12 oclock is 440hz, and the 12 chromatic notes would be mapped to the angle of the hours. Maybe red pixel intensity is bit mapped to octave, IE first, third and eight octave: 0b10100001. Time would be represented by radius. Unfortunately the space wouldn't wrap nicely like if there was a native image format for representing donuts.
- crubier 4y agoYou need the mapping to be only to 1 dimension, since you need the 2nd dimension for time
- mixeden 4y agowow!
- thedangler 4y agoAnyone know how I could try and use this with Elixir livebook with https://github.com/elixir-nx/bumblebee https://github.com/elixir-nx/bumblebee I'm new but this is something that would get me going.
- 2dvisio 4y agoVery interesting idea :) Unfortunately it breaks when I enter Tarantella or Taranta. Need more training samples from south of Italy :)
- swfsql 4y agosnake jazz
- quirkot 4y agoAnyone out there that speaks Arabic, can you let us know if the Arabic Gospel clip contains words? To a speaker do the sounds even sound Arabic?
- sansieation 4y agoThis is genuinely amazing. Like with any AI there are areas it's better at than others. I hope it doesn't go unnoticed just because people try typing "death metal" and are not happy with the results. This one seems to excel at 70-130BPM lo-fi vaporwave/washed out electronic beats. Think Ninja Tune or the modern lo-fi beats to study to. Some of this stuff genuinely sounds like something I'd encounter on Bandcamp or Soundcloud. I think I'm beginning to crack the code with this one, here's my attempt at something like a "DJ set" with this. My goal was to have the smoothest transitions possible and direct the vibe just like I would doing a regular DJ set. https://www.youtube.com/watch?v=BUBaHhDxkIc https://www.youtube.com/watch?v=BUBaHhDxkIc I wonder if this could be the future of DJing or perhaps beginning of a new trend in live music making kind of like Sonic Pi. Instead of mixing tracks together, the DJ comes up with prompts on the spot and tweaks AI parameters to achieve the desired direction. Pretty awesome.