4 ms·
This is really neat! But I think it's a stretch to call it AI-generated jazz music. As I understand it, the author has trained an LSTM on a single MIDI file --
by rryan 10y ago
This is really neat! But I think it's a stretch to call it AI-generated jazz music.
As I understand it, the author has trained an LSTM on a single MIDI file -- "And Then I Knew" by Pat Metheny. The network is then asked to generate MIDI notes in sequence.
What this network has been asked to do is to produce an output stream that is statistically similar to the single MIDI input file it has been trained on. It would be more accurate to call this an "And Then I Knew" generator. Its "cost function" -- the function the network is trying to minimize during training -- is exactly how well it reproduced the target song.
Neural networks are "universal function approximators". It's not surprising that given a single input, a network can produce outputs that are statistically similar to it.
A network that could compose novel MIDI jazz would look like this:
* Train a network on a corpus of thousands to hundreds of thousands of MIDI jazz files.
* Add significant regularization and model capacity limits to prevent the network from "memorizing" its inputs.
* Generate music somehow -- the char-RNN approach described here is fine. There are other methods.
You want the network to build representations that capture the patterns of jazz music necessary to pastiche them but not high-level enough representations that the network is exactly humming the tune "And Then I Knew". This is so much of a problem that any paper presenting a novel result in generative modeling pretty much must include a section presenting evidence their model is not memorizing its inputs.
I can hum a few classic jazz tunes from memory but that mental process is not jazz music composition -- it's reproducing something from memory. If we're going to call a model "AI-generated jazz" you need some way to tell the network to not hum a tune it knows and instead compose a new tune with the principles/patterns it knows. Since we can't speak to our models and tell them to think one way and not the other, part of the trick in this field is to come up with models that can only do one thing and not the other.
- aczerepinski 10y agoCollective improvisation is the core of jazz's identity, more so than any of its other defining traits (swing, syncopation, blues-derived harmony, etc). Generating random patterns that sound jazz-ish is interesting, but until multiple generators can react to what the other is doing in real time (or to a human participant), it isn't exactly jazz. I'd equate it to a basketball playing robot. Teaching it to shoot free throws is interesting, but doesn't really take a step towards approximating what basketball is. Can it call for picks, lead passes to cutting teammates, box out for rebounds, force bad shots, etc?
- Balgair 10y agoWell, given enough time and resources, then yes, the b-ball-bot could, and probably better than a human could. I know this is a cop-out answer, but look at the DeepMind Go games. The computer beat a top 100 (I don't know the rankings, actually) Go player, something that was thought of as nearly impossible in this decade. The most interesting thing was if you read the commentary on the matches. The announcers were mystified by the computer's moves. 'Alien' comes up a lot in describing the play-style. Us humans can't play Go and evaluate each stone in the game. We have to 'chunk' the game. Exp: These 3 stones are a 'wall' or a 'platoon', this stone is 'hot' and can take your stones, this stone is 'down' and will be used in 3 turns, etc. The computer doesn't have to do that chunking, each stone is evaluated individually. As such, the play-style was totally foreign to people. It did things no player had tried or, importantly, could have thought of given the limits of our brains having to 'chunk' the information. I would predict that a b-ball-bot would play the same way, in totally strange ways that a human can't think of. Exp: Calculating a reasonably high probability that the ball will bounce off your nose and go into the left hand of it's team-mate, throwing the ball as hard as it can at it's own head to make a shot, not trying to get past just 1 opponent but the entire team's right thighs 57 seconds from now, etc. Similarly with jazz, the computer is a dumb machine that will just do strange things because humans have to 'chunk'. In music, we play in chords and notes and with rhythm and timing. The computer can evaluate the whole song, and every other song at the same time and can borrow from all those. You and I can pull in the feelings of loss of a child, or the joy of strawberry ice-cream bars in a Memphis summer, things a computer will never. But we cannot pull in the obscure Tuvan throat singing techno-remixes on Youtube , the Afro-Thai heavy metal Vimeo channels, or the terrible pre-teen angst poems set to crappy guitar, etc, all at once. It can only see what you feed it, but you can feed it the life-outputted-into-music of billions of humans with live updates. The computer will know more. But music is emotional and about feelings. The feel of music is most important to us. And I think that a human songwriter is therefore essential, one that cares and puts effort into the work. It connects us, and that is what is important, not the sounds.
- cdr 10y ago> But music is emotional and about feelings. The feel of music is most important to us. And I think that a human songwriter is therefore essential, one that cares and puts effort into the work. It connects us, and that is what is important, not the sounds. Children can play music very emotionally (or rather, in a way that adults associate with emotional) without having any experience of or real comprehension of the emotions. Imitation and training is sufficient to be convincing. A program doesn't need to experience emotion, only know that certain characteristics of the sound are associated with certain emotions.
- Rauchg 10y agoEven then, true jazz music composition would not involve only jazz training data. Even if it's thousands or hundreds of thousands of songs. Wouldn't you just be diversifying the source for your statistical reproduction? A human composing new creative Jazz is using a much wider set of sources for creativity, not just existing jazz songs.
- rryan 10y agoI think then the question becomes, are humans doing anything different than that? If so, why do you believe the network is only reproducing statistics rather than having learned the same circuitry humans have when improvising/composing jazz? It's hard to show that it's doing one thing or the other. In this case with n=1, it seems pretty clear it's doing the former. If not, then it doesn't seem to matter since that's what humans are doing.
- ThomPete 10y agoIs it though? Most musicians learn from others and thus develop a style semi inspired by what they listen to. So add 50.000 more songs and you have something. Perception is reality.
- squeaky-clean 10y agoThis is the commenter's point. Because the current way of training the model is by comparing its own output to the Pat Metheny tune, it doesn't work once you add more than a single song. Musicians do learn from each other, but then they learn how to play what they like, or what sounds good. To this model, "what it likes" is a 100% representation of 'And Then I Knew' . You could swap the target song for another, but not for multiple targets at the same time without reworking the logic.
- ThomPete 10y agoIsn't that just a matter of expanding the feedback loop to include other things than just the music? I.e. aren't the mechanics there and is primarily limited by the the size of the feedback loop? I am genuinely interested in the answer.
- jlu 10y ago@rryan mind sharing methods other than char-RNN? Thanks!