6 ms·
How to Train a Gen AI Kick Drum Model on Your Old Linux Desktop with 6GB VRAM
- Capitanai 3mo ago[flagged]
- pringk02 3mo agoI just wish it had samples! I want to hear it
- zhinit 3mo agoFor sure! I just added a couple
- motoxpro 3mo agoThey sound cool! Add a few more! :P
- jdalgetty 3mo agoI was hoping to hear some songs using these samples!
- andai 3mo agoDid you save any of the "failed" results? I'd love to hear what kind of weird sounds it makes out of distribution (e.g. on the keywords it didn't have much data for).
- zhinit 3mo agoI just added one with the "techno" keyword in the keywords section. Its pretty weird lol
- scoot 3mo agoNot sure if there's something wrong with the player, or if it's just me, but they both sound like noise. I guess the first sounds vaguely kick drum-like (but distorted), the second is just noise. Chrome 149.0.7827.200 (Official Build) (arm64), macOS Tahoe 26.0.1
- kleiba2 3mo agoI have to admit I don't understand what exactly the problem is we're trying to solve with ML here...?
- tgv 3mo agoDecomposing sounds from (fully produced?) tracks into underlying components, and then giving the user the option to synthesize them with different parameter settings. I think.
- pennomi 3mo agoI was trying (and failing) to do this the other day. It’s a really interesting problem space and I love to see someone with a more solid foundation give it a try.
- tgv 3mo agoThere is better software, although it might not meet OP's goal. There's e.g. Melodyine, which can identify (and change) notes from individual instruments in polyphonic passages, and there are also tools to estimate and remove reverb and other forms of processing. Those are based on classical DSP. OP just wanted to use "AI".
- zhinit 3mo agoI wouldn't exactly say it's trying to solve a problem. It's to explore and see what happens which is what music is all about. It's also a unique niche model I haven't seen before.
- mock-possum 3mo agoThis is a really really fun sounding project - ironically, because there are no audio samples provided at all. I would have thought a music producer creating samples for music would naturally let you listen to what they were making.
- zhinit 3mo agoGood point! I just added a couple at the end of the intro
- deleted 3mo ago[deleted]
- lardosaurusrex 3mo agoI always roll my eyes when I see LLM weirdos talk about getting models to run on "old" hardware and finding out it's hardware that's still better than what most people have access to. It doesn't make it any less impressive to those who know what hardware requirements for LLMs usually is/are but for those with no idea it usually ends up reinforcing bitterness towards it as they feel annoyed that their own hardware is somehow worse and yet are unable to upgrade because of said LLMs stealing all the hardware in the world all while RAM/memory/storage manufacturers manipulate the market(s) against them.
- zhinit 3mo agoFor sure. If you are curious I used a NVIDIA GeForce GTX 1660 SUPER So to be exact, it came out 7 years ago (I upgraded at some point on this desktop a long time ago and didn't remember the exact year) (I updated the article to reflect this now) This cost $230 new and you can get one now for $100 which I don't think is too out of reach.
- jdboyd 3mo agoThe Geforce GTX 1060 launched 10 years ago with a MSRP of $249. It spend 5 years and 4 months as the #1 card according to the Steam Hardware Survey. That makes it hard to feel that it is fair to accuse it of still being better than what most people have access to, unless you are asserting that most people have access to no GPU at all, which is likely accurate, but not likely to be accurate here, nor in any sort of enthusiast circumstance. If you lump the Intel Xe built in graphics (started with the 11th gen Core Is) and the Intel UHD (launched with 8th gen Core Is) together, the combined group would come in 6th place, with the 5 places above that in commonness for people who are actively playing steam games all being considerably faster than the Geforce GTX 1060 or Geforce GTX 1660 cards. Interestingly, now the #1 GPU is the GeForce RTX 4060 Mobile version, which I believe is the first time the top has been a laptop chip instead of desktop chip.Items #2 and #3 on the list are the 2 generation old RTX 3060, followed by the 1 generation newer RTX 4060. 4th and 5th are RTX 5070 and RTX 3050.
- juancn 3mo agoExcellent article! I think it has the right level of detail, one question though: why the shape of the tensor? 4x8x11. That I didn't get from the text.
- zhinit 3mo agothe spectrograms are 128x173 (128 mel frequency bins by 173 time frames) the encoder is downsampling 4 stages of stride 2 convolutions so it halves dimensions 4 times 0: 128 x 173 1: 64 x 87 2: 32 x 44 3: 16 x 22 4: 8 x 11 Then i used 4 separate channels. This was somewhat arbitrary due to the local training constraint. This would be a hyper parameter worth tuning if I had time to dig into this more. I trained this a few month ago and don't remember exactly what I tried before I arrived here, but I only ran the whole process 2 or 3 times because of how long it took to train. Hope this answers your question!
- juancn 3mo agoYeah, thanks!!
- johndear223 3mo agoArticles like this are why I come back to HN. Interesting technically, kinda novel and fun. Got me thinking about datasets that may be sitting on old HDD, got TBs of old video and audio from projects of past. Blogs like this help point the way.. Now if only I had the time..
- dj_axl 3mo agoModeled reverb yet no modeled compressor, hrmm, is compression not used on kick drums (or not a big part of the sound)?
- zhinit 3mo agoThe compression is the OTT which stands for Over The Top compression. It was originally a multiband compressor preset in ableton and is now used widely throughout dance music.
- larme 3mo agoPeople who are interested in this application should check synplant[0]. It has a ML technology called "Genopatch" which gives you 2 functionality: 1. you can try to describe a sound with some tags and it will try to generate a sound to capture the feeling of these tags 2.you can feed it with a sound sample and it will try to re-synthesize the sound with its synth engine. Though the end result will usually be just a "re-imagined" version of your input sample. My guess is the underlying model is not a "deep" model. The main benefit is that the end result is not a wave file, but a list of generated parameters that can be synthesized by the synthplant engine. And now it comes the interesting part: you can tweak these parameters to finetune the generated sound. These parameters have actual meanings (FM ratio, reverb etc.) [0]: https://soniccharge.com/synplant https://soniccharge.com/synplant
- aw123 3mo agoHow far are we from getting a general model that can resynthesize any instrumental audio sound without fiddling with any knobs, so that we can recreate instruments we hear from any song? Seems like it should exist by now?
- zhinit 3mo agoSUNO is pretty close. It still has some weird things going on with high frequency artifacts and phase between left and right channels but if you aren't listening on a good system (like a phone) most people probably wont notice.
- larme 3mo agoFor me creating the exact sound is not very interesting from sound designing perspective. You can always sample the real instrument. Like physical modeling synthesis, the interesting part is to compress the sound to some parameters that you can tweak and generate new sounds Another approach is VAE, which also you give your some latent embedding, you can tweak the embedding to generate new sound. However the meaning of this embedding is not explicit.
- lightedman 3mo ago
- trencedamp 3mo agoFor a moment I thought Gen AI meant the current generation of kids. It's a fitting moniker
- cocodill 3mo agosomeone needs to take care of the snares
- zhinit 3mo agoIf you are committed the model should work about the same on any type of one shot sample. The code is public and documented so if you have the snare collection and a macbook you could probably point claude/chatgpt at it and it would be able to train on your laptop.
- tunk 3mo agoThis has been done years ago. See https://audialab.com/products/emergent-drums-2/ https://audialab.com/products/emergent-drums-2/ for instance.
- zhinit 3mo agoInteresting! I had not seen this. On their website they mention diffusion but not the other models so it might not be identical but its definitely similar.
- tunk 3mo agoThat's fair :) I also could've phrased my comment a bit more politely.
- drcongo 3mo agoIt reminded me of the (possibly apocryphal) story about Liam Gallagher trashing a hotel room on tour - when asked by a reporter "why? It's been done before." he supposedly replied "yeah, but not by me". Sometimes the "by me" is the interesting / fun / instructive part.
- baerbelblue 3mo ago[dead]
- robotswantdata 3mo agoConfused. Why not just make the kick drum from a sine? Seconds
- zhinit 3mo agoI sound design a lot of stuff (in fact I made some of the default kicks in the app), but this is just a different tool, and I wanted some practice training and deploying a generative AI model.
- lilbigdoot 3mo agoA sine is only a part of most kicks
- robotswantdata 3mo agoEverything is a sine
- defrost 3mo agoco-signed.
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- psd1 3mo agoThank you, thank you, we're the Fast Fouriers, good night!
- deleted 3mo ago[deleted]
- Cloudef 3mo agoIts possible and boy do i have a youtube channel for you https://youtu.be/ndG-6-vONNc https://youtu.be/ndG-6-vONNc https://youtu.be/8dfgum9XlJc https://youtu.be/8dfgum9XlJc
- thangalin 3mo agoSlightly off-topic. Now that 1920s jazz music is falling into public domain, has anyone tried to reinvigorate the music using AI and generative adversarial approaches? Pre-1940s music didn't have high-fidelity sound, so the strong bass lines weren't captured. In theory, we could "downgrade" modern recordings to sound like 1920s recordings, then use adversarial techniques to train the machine on how to restore the antique recordings. Anyone know of any work being done in this area?
- gregdaniels421 3mo agoSo the idea would be to reconstruct the low frequency components from whatever upper harmonics are left in the recording? If you know the instruments and positioning of the recording device and something of its(the instruments, recorder, environment, etc.) characteristics, it might be possible to solve that using classic methods. There would be huge numbers of parameters, it is an interesting thought. Is there a large easily/freely available corpus of those recordings?
- zhinit 3mo agoTo do this, I think you are right that you would need to 'downgrade' modern recordings to sound old so that you have both sides of the training data covered. This would be a cool project to work on. Ideally you would buy some vintage gear and then run the audio through both, but that would be very expensive. You could may be find some vst emulations though and get decent results.
- wildzzz 3mo agoIt might be easier than that. Are the bass lines totally missing or are they just very weak? If you can capture a recording using vintage equipment and the placing of it, you can get the system response. Run the original recordings through an inversion of the response and you should get really close. Another possible method is to find the transform between an identical modern recording of the song and use the difference between the two recordings to make your transform.
- embedding-shape 3mo ago
- fabiofzero 3mo agoYou could save so much time and processing power by just learning how to sidechain.
- bgenchel 3mo agonice work!
- zhinit 3mo agoThanks!
- BrandNewRetro 3mo agoI have three (and pray I do not come up with any more) needs for AI audio apps. One, a ubiquitous restoration model. Find degraded copies of music in the wild, old YouTube's, transcodes, vinyl rips, bad masters, half destroyed tapes... Pair them with modern pristine lossless encodes of the same music, train. Then use that model on music we don't have pristine copies of. The second is similar but more specific. There are so many stems floating around from popular music. My idea is to compare individual stems against the results from MVS/Spleeter(same song, same instrument). This would surely stand a chance of pushing that tech forward, so we can treat the FFT artefact heavy sound of new efforts. Thirdly, from a creative point of view, I wanna do the equivalent of image to image on my tracks... But I actually want it to hallucinate in the manner of the early deep dream images, I want to be able to play with that space.. I can knock out musak to spek in minutes already, gen music is just reducing low effort to nearly no effort, preventing people with needs from networking with creators.... Uhh.. but I think that's a very general issue with Gen AI away from the corporate/entrepreneurial dev space
- zhinit 3mo agoFor 1) another commenter had a simmilar idea. I think a GAN would be able to do this. It's just a matter of collecting the data For 2) SUNO's stem separation is pretty good. The open sources ones (like spleeter) are also not bad, but they are pretty hit or miss. For 3) This is a really cool idea. The deep dream images produced the fur and eyeball textures because they used imagenet data which had a bunch of dogs and other animals in it. The trick would be finding a good data set to use for this. May be there is a public source of animal sounds floating around you could try this with. And yeah think people 100% generating songs is pretty wack, but I could see why it would be fun to someone who is not a musician.
- Tade0 3mo ago> A latent space is just a compressed representation of something. I love these one-liner explanations - the absolute minimum information one needs to move forward. Great writing.
- farlow 3mo agoGreat writing and cool project! I really enjoyed your intuitive explanations.
- zhinit 3mo agoThanks!
- mk_stjames 3mo agoThis is cool, I had some questions because I was trying to understand this more - not the diffusion model training part, because that I've actually seen before- but the actual way you bundled this into a web application and the choices made... Did you publish the actual trained model anywhere? I see how in the code there is python for how individual samples can be generated, but the model checkpoint pulldown comes from a directory that... I don't see. I then went through the code of how this runs on the web and- I'm not a web dev guy- so I'm pretty confused at all the bits bolted together to make this into a web app. It seems like there is a WASM bit that is compiled from a typical C++ audio plugin that is doing the stuff like conv reverb and limiting and distortion in the web app - all that is oldschool, non-generative AI, DSP being applied to the samples. Then the samples are just... a few default generated samples, to start- where are they pulled from, physically? And you have a login requirement to spool up the actual generative AI part to generate new samples to run into the DSP (because that needs a GPU on the backend to do, so, a login to help rate limit this) How big is the actual generative model? Did you ever think about building the generation engine into the WASM bundle, using maybe WebGPU in the WASM to accelerate in a platform agnostic way, so that the entire app would run offline in someone's browser window? I'm having fun just playing with the kick program without a login, which, again, am I right in saying in that mode there is no gen AI of sampled happening server side, it is just playing starting with some pre-made samples?
- zhinit 3mo agoThe weights are ~300MB and are on hugging face here. https://huggingface.co/zhinit/kick-gen-v1 https://huggingface.co/zhinit/kick-gen-v1 It was a lot of work to get a good DSP to work on the web hahah. Yes, I am writing the DSP in C++ and compiling to WASM. Im using multithreading so the audio work is done in the AudioWorklet while the UI is run in the main thread. I was thinking of writing another article on this because it's pretty interesting and a bit complex. I sound designed some of the kicks myself and some of them are from sample packs. I just renamed them all to have German names. If I wanted the model to run completely in the frontend I would need the user to download 300MB of weights and it would probably be tricky making sure it works on everyone's hardware. So I though about this but it didn't seem like the best option. I'm pretty sure it is possible though. And yes I put rate limiting so no one goes crazy on my credits. Im glad you're enjoying it!
- codedokode 3mo agoWhat about copyright? Can one take 13,000 copyrighted works, train the model and release it for free so that nobody ever buys any of those 13,000 samples? Can one (legally) convert all commercial sample libraries into a free neural network?
- zhinit 3mo agoIm not selling a product here so I'm not going to worry about it for my purposes. This is the hot topic in AI ethics. Is it just learning the same way that a human learns from hearing songs on the radio and playing them, or is it just a compression algorithm? There is not a clear answer here but it seems like the legal system so far is letting things slide and agreeing with the former argument. (This can also be applied to other areas like open AI reading the NYT)
- nativeit 3mo agoI remember making kick drum models with kilobytes of RAM.
- thenthenthen 3mo agoCouldnt get sound to work on iOS but back at my Desktop now and wow these examples/presets sound really cool, great job! (The examples in the write up sound very weird btw...). Try out the 'app' here > https://kick-with-reverb.vercel.app/ https://kick-with-reverb.vercel.app/
- zhinit 3mo agoYeah, for web apps you have to make sure your phone is not on silent and the ringer volume is up. This is why most audio apps for phones will force you to use their native app. I was looking into a way around this and I don't think it exists without building out an entire phone app.
- seth17 3mo agoHave you tried setting `navigator.audioSession.type = playback`?
- zhinit 3mo agoThank you! This one liner worked. I searched for a solution for this when I was building this app and am not sure why I couldn't find this. Looks like this has been possible since iOS 17 but is not well documented.
- thenthenthen 3mo agoIm building some web audio things and iOS is really a pain. Sometimes the audio plays, sometimes not, sometimes it get stuck playing in the bg. The silent mode button is also something i forget often haha thanks for reminding me again (my phone is always on silent because i get about 15 spam calls a day…)