5 ms·
The analysis was interesting but the end result was pretty terrible. Of course music in the past was often considered terrible by the following generation, a tr
by coldcode 5y ago
The analysis was interesting but the end result was pretty terrible. Of course music in the past was often considered terrible by the following generation, a trend that still exists today.
What I would really like to see is an attempt to write multipart contrapuntal works like Bach with some kind of AI. The rules are fairly well understood, but Bach knew how to adapt and even violate them all but still wind up with amazing pleasant music.
- rsfern 5y agoMeh. There’s plenty of deep learning music generation stuff out there, this is still a really cool approach I think it would be cool to combine the two. Instead of generating raw midi, your GAN or reinforcement learning agent or whatever could try to generate sequences of transformations to melodic fragments. Neural program synthesis type stuff. Or maybe one could build an automatic music analysis tool that can start from the score and try to infer the program that generated them. (Is that a thing already?)
- TheOtherHobbes 5y agoThe reason the end result is pretty terrible is because classical harmony and melody are closely related, and you can't job-lot-replace one without ruining the other. This project is quite similar to something I've been working on, but I realised early on that you can't split up features like this and get credible results - credible meaning "appropriate for the style grammar." It's a bit - only a bit, but let's go with it - like trying to generate sentences by swapping out nouns and adverbs. You end up with something that is grammatically correct in theory but makes no sense in practice. Classical music particularly is fundamentally integrated in a way that textbook analysis doesn't fully explore.
- zozbot234 5y agoThis comment should not be disregarded so easily. The reason why deep sequence learning has the best results in generating complex, highly contrapuntal music (it's more like noodling or improvisation than an actual compositional process, but it is generally compelling at its best) is precisely because of the loosely grammar-like structure mentioned in OP. The algorithmic operations they play with are not very well defined but the background theory is sound, and closely reflects what music theorists and composers in general have written about the subject in the 500 years or more it has been seriously studied. As for deep learning models which create good contrapuntal music, see e.g. 'Biaxial RNN' https://github.com/danieldjohnson/biaxial-rnn-music-composition https://github.com/danieldjohnson/biaxial-rnn-music-composit... by Daniel D. Johnson, who is now at Google Brain but wrote this as an independent(!) researcher. (Note that the existing code requires Python 2.x It would be interesting to forward-port it so it can work with Python 3.x and a maintained version of Theano. Replicating the model using Tensorflow would also be quite worthwhile.) If you're interested in Bach's work specifically, the "BachBot" and "DeepBach" projects are also interesting but less accessible. Example output for all of these models can be found on the Internet, just look around for it. The proprietary system AIVA is also worth mentioning because even though it's so proprietary and secretive, the compelling and "serendipitous" music it manages to come up with is a tell-tale sign that it's actually doing well-founded deep learning stuff behind the scenes, much like the aforementioned open systems. Note that much of the released output has been orchestrated (AFAICT) manually by humans, but at some point I was able to find some piano-format reductions that are most likely very close to what the AI actually created, somewhere on the official site.
- p1esk 5y agoDeep learning needs a lot if training data. There’s not enough MIDI encoded music to train something like a GPT-3 model to generate MIDI sequences. A better way is to train on and generate raw audio, see OpenAI JukeBox. Unfortunately it’s extremely compute intensive (even compared to GPT-3), so it will probably be a few more years until they (or some other big player) releases JukeBox-2.
- zozbot234 5y agoThe amount of data you need depends on the model architecture you're using. A generic model like GPT-3 is neither here nor there, but something specifically intended for music can make do with very little data.
- p1esk 5y agoWhat do you mean “neither here nor there”?
- zozbot234 5y agoThe structure of something like GPT-3 is far too weak and general to achieve good results for something as structurally complex as music. It's designed to generate text - and then mostly natural language text. Music is very different, as OP hints in the linked post.
- p1esk 5y agoI thought you wanted to use deep learning. In DL transformers are the best we got currently. Both JukeBox and MuseNet use transformers. What makes you think they are not up to the task?
- zozbot234 5y ago"Transformers" is a general technique, comparable to "LSTM" or "attention". The details of just how much prior information about the domain you reflect in the model architecture are just as relevant, and it's not clear just how well MuseNet addresses this compared to the earlier work I mentioned above. As for JukeBox, it's trying to solve audio generation at the raw samples level, which is quite literally a complexity increase of several orders of magnitude.
- Rochus 5y agoSee e.g. https://github.com/feynmanliang/bachbot https://github.com/feynmanliang/bachbot. The companion site is no longer available, but here are some results on soundcloud: https://soundcloud.com/bachbot https://soundcloud.com/bachbot Here is another very good one: https://openai.com/blog/musenet/ https://openai.com/blog/musenet/ I'm a trained musician myself and interested in automatic music composition following the progress for the last thirty years, but only recent work (like the ones referenced) produce convincing results (besides Cope's work of course, but which required manual selection and editing). You might also be interested in this survey paper: https://arxiv.org/abs/1709.01620 https://arxiv.org/abs/1709.01620
- ekianjo 5y agosid meier had a software like that he created for the 3do
- hackoo 5y agoIs "the end result" the Beethoven's sonata in Japanese scale? If so, the most important reason that it sounds terrible is the music is generated with MuseScore, without adjusting dynamics, tempos, etc. Actually, the Beethoven's original sonata in this blog is also generated with MuseScore, and it sounds not so good even with dynamics added. However, this is why I agree that deep leaning is more promising than this manual approach, since too many variables you need to adjust to make music sound good rather than syntactically correct.
- whiddershins 5y agoI don’t agree at all, and it’s not even experimental. And I think that is not a very nice thing to say about other people’s work.
- klyrs 5y ago"Pretty terrible" is awfully strong. I've heard a lot of terrible music and this doesn't make that cut. What you're reading is a nice writeup of some theory, with a worked example to showcase how to utilize some features of a library which is a work in progress. The author even notes that the music is simplified for the example. It isn't meant to be Bethoven. It isn't even AI. If you want to see AI, other people have done that already. What's cool about this library isn't the quality of the (midi, retch) output -- what's cool is the actual Python library behind it, and the writeup, which are both super easy to read and follow.
- Asooka 5y agoI thought it sounded really good!