6 ms·
AI is bad at music also. Even the state of the art transformer models can't produce more than a few seconds of coherent melodic phrases.
by dimmuborgir 4y ago
AI is bad at music also. Even the state of the art transformer models can't produce more than a few seconds of coherent melodic phrases.
- CactusOnFire 4y agoAI is bad at Audio. AI can do MIDI fine.
- causi 4y agoWhich is a real shame. AI-powered restoration of poor-quality audio would be highly useful.
- aaroninsf 4y agoThat particular niche has had some pretty amazing successes already. It's coming. We can't produce arbitrary media streams with many "stack layers" of meaning and detail yet, but we can do a lot of specific instrumental transformations... Vaguely relevant: https://koe.ai/recast/ https://koe.ai/recast/
- dwringer 4y agoMIDI is extraordinarily expressive and is likely used to sequence a large majority of music produced within the last three decades. A lot of the instruments you hear are synthesizers or samplers running directly from MIDI. There is a lot more to what MIDI can do, and is used for, than the conception most people have from "canyon.mid" or old website background music. If an AI can do MIDI just fine then it's an extremely small leap to doing audio just fine.
- p1esk 4y agoIf an AI can do MIDI just fine then it's an extremely small leap to doing audio just fine. Unfortunately this is not true. It takes a huge amount of human effort to make MIDI encoded music sound good. The difference between MIDI and raw audio music generation is the same as the difference between drawing a cartoon and producing a photograph. To clarify, yes MIDI can be expressive, but what's being generated when people say "AI generates MIDI music" is basically a piano roll.
- NateEag 4y agothank you for clarifying this. As a clasically-trained pianist who then got into electronica and synthesis, it was mind blowing to me that people could wrangle expression and phrasing from a MIDI sequencer.
- dwringer 4y agoI'm not familiar enough with existing implementations of such systems to dispute it, but there's no fundamental reason algorithmic composition systems could not include modulation parameters of all kinds (pitch/breath/effects/synthesizer controls/etc) in their output. I am envisioning a DAW set up with several VST's and samplers with routing and effects in place, then using some combination of genetic algorithms and other methods to "tweak the knobs" in the search for something pleasing. The search space is absolutely enormous, though, so I don't dispute that it's very difficult, but I wouldn't go so far as to say that it can't be done. In such a space there are "no wrong answers" so to speak. I have a python script which creates randomized sequences of notes/rhythm and gives each one a different combination of LP/HP filters and random envelopes - it's not music but it takes on a much less mechanical quality by emulating different attacks and timbres over time, even though it's completely random. I would go so far as to say I'd be genuinely surprised if algorithmic composition and production hasn't been used to some extent significantly greater than "basically a piano roll" in at least some of the past decade's top 40 music on the radio.
- p1esk 4y agothere's no fundamental reason algorithmic composition systems could not include modulation parameters of all kinds (pitch/breath/effects/synthesizer controls/etc) in their output There is such a reason - lack of training data. Very few high quality detailed MIDI samples exist to train machine learning models like AudioLM. For state of the art in MIDI generation, take a look at what https://aiva.ai/ https://aiva.ai/ produces (it's free for personal use). There you can compare raw MIDI output to an automatically generated mp3 output (using "VST's and samplers with routing and effects in place, then using some combination of genetic algorithms and other methods to "tweak the knobs" in the search for something pleasing.") mp3 version will sound much better than raw MIDI, but (usually) significantly worse than music recorded in a studio and arranged/processed by a human.
- stephencanon 4y agoWhich is extra funny, because GOFAI models (e.g. David Cope's work) were doing a pretty OK job back in the 1990s!
- mjburgess 4y agoI think if we replaced "AI" with "taking averages over subsets of historical examples", then there'd be no mystery for when "AI" will be good or bad at anything. Would we expect a discrete melodic structure to be expressible as averages of prior music? No.
- yeasurebut 4y agoThat’s what a musician does. They make short loops and loop them. This reads like someone who knows sheet music and theory but does not listen to music. It’s repetition of short phrases over and over. I’m not really sure what people expect of general AI trained on human generated outputs. It can’t make up anything anything “net new” only compose based upon what we feed it. I like to think AI is just showing us how simple minded we really are and how our habit of sharing vain fairy tales about history makes us believe we’re masters of the universe.
- dimmuborgir 4y agoThose models are not trained on short loops. They are trained on whole songs just like image generation models are trained on whole images. And yet they struggle to repeat sections, modulate to a different key, create bridges, intros and outros. After a few seconds of hallucinating a melodic line they simply abandon the idea and migrate to another one. There is no global structure whatsoever.
- yeasurebut 4y agoMusicians don’t spit out an album in one sitting and they’re highly trained in theory. They get bored and tired of a process and take breaks. They come up with an album of loops composed together over time. AIs state will forever be constrained to the limits of human cognition and behavior as that’s what it’s trained on. I read published research all year. Circular reasoning. Tautology. It’s all over PhD thesis. There’s no “global structure” to humanity. Relativity is a bitch. Seeing the world through the vacuum of embedded inner monologue ignores the constraints of the physical one. It’s exhausting dealing with the mentality some clean room idea we imagine in a hammock can actually exist in a universe being ripped asunder by entropy. It’s living in memory of what we were sold; some ideal state. Very akin to religious and nation state idealism.
- mjburgess 4y agoI think it's deeply depressing that AI has been sold as something even capable of modelling anything humans do; and quite depressing that this comment exists. "AI" is just taking `mean()` over our choice of encodings of our choice of measurements of our selection of things we've created. There is as much "alike humans" in patterns in tree bark. AI is an embarrassingly dumb procedure, incapable of the most basic homology with anything any animal has ever done; us especially. We are embedded in our environments, on which we act, and which act on us. In doing so we physically grow, mould our structure and that of our environment, and develop sensory-motor conceptualisations of the world. Everything we do, every act of the imagination or of movement of our limbs, is preconditioned-on and symptomatic-of our profound understanding of the world and how we are in it. The idea that `mean(424,34324,223123,3424,....)` even has any revelance to us at all is quite absurd. The idea that such a thing might sound pleasant thru' a speaker, irrelevant. This is a product of i dont know what. On the optimist side, a cultish desire to see Science produce a new utopia. On the pessimisst side, a likewise delusional desire to see Humans as dumb machines. What a sad state!
- vladf 4y agoHave you heard the piano continuations of AudioLM? https://google-research.github.io/seanet/audiolm/examples/ https://google-research.github.io/seanet/audiolm/examples/
- bloep 4y agoIndeed, there is lots of denial or ignorance in this thread (ignorance in the technical sense). AudioLM already produced impressive results and it's a tiny fraction of what is already possible because performance simply improves with scale. One can probably solve music generation today with a ~$1B budget for most purposes like film or game music, or personalized soundtracks. This is not science fiction.
- p1esk 4y agoI don't see a lot of progress in AudioLM compared to results from 2018: https://storage.googleapis.com/magentadata/papers/maestro/index.html https://storage.googleapis.com/magentadata/papers/maestro/in... What's more interesting and concerning - listen carefully to the first piano continuation example from AudioLM, notice the similarity of the last 7 seconds to Moonlight sonata: https://youtu.be/4Tr0otuiQuU?t=516 https://youtu.be/4Tr0otuiQuU?t=516 I'm afraid we will see a lot of this with music generation models in the near future.
- bloep 4y agoThere are quite simple tricks to avoid repetition/copying in NNs, e.g. by (1) training a model to predict the "popularity" of the main model's outputs and penalizing popular/copied productions by backpropping through that model so as to decrease the predicted popularity, or (2) by conditioning on random inputs (LLMs can be prompted with imaginary "ID XXX" prefixes before each example to mitigate repetitions), or (3) by increasing temperature or optimizing for higher entropy. LLM outputs are already extremely diverse and verbatim copying is not a huge issue at all. The point being, all evidence points to this not being a show stopper if you massage these evolutionary methods for long enough in one or more of the various right ways.
- deleted 4y ago[deleted]
- aaroninsf 4y agoAI can be quite good at music, but yes there is not yet at on-demand button rendering from a text prompt of bitstreams encoding composed performed and mastered music.
- denton-scratch 4y agoIt doesn't surprise me that an AI model for language can't grok maths or music. I can't see how a language model can map to maths. Hell, I don't even know how to describe music in words. It's possible to articulate some maths in words, but that often involves using words with unexpected definitions.
- Der_Einzige 4y agoThat's wrong, and shows how ignorant you are of SOTA techniques for music generation. They are far ahead of that.