5 ms·
The author mentions about 5% of the generated music is any good to take a listen. Even the best sounding ones shared don't sound that good to me.
by mebr 6y ago
The author mentions about 5% of the generated music is any good to take a listen. Even the best sounding ones shared don't sound that good to me.
- nr2x 6y agoWent in with high hopes, but it’s not remotely listenable.
- narag 6y agoMost of it sounds terrible. I guess that it's because there is no way to give some feedback. A chess program knows what means winning. And IIRC "this person doesn't exist" uses human opinion to train. Music has many rules, not only theoretical but unwritten rules about what works. You must incorporate them somehow into the program, either by code or offering something to the program to deduce them.
- deleted 6y ago[deleted]
- gwern 6y agoI previously tried an approach which uses DRL for feedback but I couldn't quite get it to work: https://www.gwern.net/GPT-2-preference-learning https://www.gwern.net/GPT-2-preference-learning At the moment, it probably would be more practical to train a model to predict ratings and use that to screen generated samples or possible completions and throw out too-low-scoring ones (the 'ranker' approach worked out very well for the Meena chatbot recently: https://arxiv.org/abs/2001.09977 https://arxiv.org/abs/2001.09977 )
- p1esk 6y agoWhat do you think about training a binary classifier to distinguish between human and generated samples? E.g. choose a state of the art model designed to classify composers or styles, and finetune it for this task.
- mattkrause 6y agoThat's a tiny step away from a GAN, which works very well for image generation, so seems promising here too.
- p1esk 6y agoGANs have been tried many times already for music generation, without much success. GPT-2 works very well for text generation so it seemed promising here too. Music falls somewhere in between text (as a sequence of chords or PCM samples) and image (as a piano roll or a spectrogram), so maybe some hybrid of image and text generators is needed.
- TheOtherHobbes 6y agoMusic does neither, which is why these naive approaches don't work.
- gwern 6y agoI think that could work potentially as a rejection sampling, but you also have the risk that it will simply find some small discriminative detail and be unusefully good at the classification; that's why you do it in a loop as a GAN, but as you mention, GANs work really badly on sequences, still, so... If you wanted to improve my ABC-MIDI GPT-2, the most straightforward ways would be to do data cleaning (I'm sure there's tens of thousands of awful MIDI files which should be removed! data cleaning with RNNs or GPT-2 or GANs always makes a large difference) and increase the model size (the fact that loss bottomed at 0.20, which is still quite bad, suggests that MIDI is hard enough that GPT-2 is struggling). More interesting would be to use Reformer or another long-range Transformer and try to operate directly on a more raw representation, like the the piano roll representation of MIDI. I think GPT-2 makes a lot of syntax errors which cripple outputs when a 'voice' goes silent, and a piano roll representation would be a lot more robust (at the cost of being like 10x larger).
- p1esk 6y agoHow would you present a piano roll to a transformer (e.g. what would be a sample of the sequence)? You could try using a tuple of pitch integers for each time step. I'm not sure how big a "vocabulary" would need to be to capture most of the chords (note combinations) - it might actually be comparable in size to a language vocabulary (tens of thousands of words). You could use two channels to capture note onset/offset info (like it was done in a biaxial RNN paper). Or the encoding used for Musenet (with explicit timing info), but somehow I like the idea of "chords as words" better.
- deleted 6y ago[deleted]
- gwern 6y ago> And IIRC "this person doesn't exist" uses human opinion to train. It doesn't, incidentally. It's just a standard StyleGAN dumping random images, trained to to model the average image/distribution, and not optimizing for human ratings or anything. Almost all the 'X Does Not Exist' things operate that way, including my own https://www.thiswaifudoesnotexist.net/ https://www.thiswaifudoesnotexist.net/
- meatsock 6y agothe rate among humans who are learning to compose is lower as ideas you don't end up using are silent.
- sq_ 6y agoI thought it was interesting, even if the overall state of most of them wasn't great. Some of them had decent sections, but then there would be some periodically repeated jarring note. Definitely better than I would've expected.
- themodelplumber 6y agoThat aspect reminds me of previous experiences with fractal art. If you're evaluating generated sets, there's always going to be a part of an original piece that's off, that's for sure. "Hey, it looks like headlights of cars driving around in the fog! But this other part makes zero sense and is distracting." If you can use a piece as inspiration, browsing around until you find the 5% or whatever % that you like, there's really some benefit to be found. I felt a bit like Simon Cowell, auditioning works and being really picky, last time I did this. But in the end you can discard an item, go with it as-is, adapt it somehow, or use it as reference material. Eventually you build a gallery.
- lostcolony 6y agoYeah...I'm now mildly curious what could be done using this, plus a genetic algorithm decided either by a human listener, or a suitability function encoding basic ideas in music.
- randyrand 6y ago"LMD: catchy ambient-esque piano piece" is pretty good. I could see it as background music in a game. I like "Pop MIDI, rapid jazz piano?" too.
- kybernetikos 6y agoYou're right. I listened to a few and didn't like them much but when I saw your comment I went back to that one, it's pretty good. It could do with a related but different section in the middle, but apart from that I enjoyed it. Do people ever train at different hierarchical levels? I've done very little ml, but it seems to me that it'd be beneficial to train a net on "plans" and then separately train one to interpret plans.
- craze3 6y agoBare in mind that these exact same MIDI notes can be fed into better sounding instruments (such as lush pads, sharp lead synths, acoustic guitars, etc...), thus improving the overall end product.
- cthor 6y agoYeah. I think a lot of people are ignoring that the output target, MIDI, is pretty limited. A skilled producer could pretty easily take these exact notes and make a track that sounds great. Sound design makes a huge difference.
- tartoran 6y agoAnd the constant velocity doesn’t help either. Almost any MIDI score could be made sounding twice as good with a little variation in velocity and slight timing errors. Speaking of which, some friends who had a video production company were picking an audio track from a commercial DVD package and there were song choice in around few dozen different styles to pick from: country, progressive rock, etc. At some point I realized that the same score that was played within different style and textures and different cadence made it sound like different songs altogehter.