4 ms·
Maybe someone more knowledgable with music theory can chime in, but the generated tunes sounds off to me. Bit like a render in the uncanny valley. Something is
by zyang 3y ago
Maybe someone more knowledgable with music theory can chime in, but the generated tunes sounds off to me. Bit like a render in the uncanny valley. Something is wrong but I can't put my finger on it.
- kweingar 3y agoYeah, it’s basically stable diffusion but for music
- drummojg 3y agoI used "rollicking" in one description and it was exactly what it sounds like to your unconscious when you are far too drunk at a country bar
- famouswaffles 3y agoNot Stable Diffusion. MusicLM is a neural codec language model. Basically if GPT was predicting the next audio token while conditioned on text.
- ShamelessC 3y agoSo like DALLE-1? FWIW, the various diffusion models (tend to) use cross attention to an attention-based approach.
- peterkos 3y agoThese models aren't actually doing any musical comparison -- they are trained on audio, and from audio, piece apart with a "note" and a "melody" and an "instrument" are from the labelled training data. No intentional theory is being done! Algorithmic music composition has usually been split into two: 1. Generate notes (re: theory, genre) 2. Generate sound (i.e., EMI[0], Kulitta[1], MusicNet[2]) Now we are doing both at the same time, and backwards. The model isn't (necessarily) going "write melody, then generate the sound", but rather, "here are 500 songs that are described with X, 500 with Y, and you want XY, so we'll combine these two" :) (This is my best understanding, so feel free to correct) [0]: http://artsites.ucsc.edu/faculty/cope/experiments.htm http://artsites.ucsc.edu/faculty/cope/experiments.htm [1]: https://hackage.haskell.org/package/Kulitta https://hackage.haskell.org/package/Kulitta [2]: https://zenodo.org/record/5120004 https://zenodo.org/record/5120004
- TheOtherHobbes 3y agoThat is it, pretty much. The problem is that coherent musical structures are much more constrained. You can't just XYZ... into a space and get something that makes sense. That will kind of work for low-density music, which includes a lot of landfill dance + subgenres. But these statistical models are blind to larger and more complex structures, and completely unaware of cultural context and semantics. It's actually a harder problem than language modelling because the spaces and the grammars are much larger, especially once you start including sound quality and production values as well as arrangement and core composition.
- laratied 3y ago[dead]
- zarzavat 3y agoThis strikes me as the wrong approach. What is the end goal here? To have an AI black box that spits out an infinite stream of music? I don’t think people are going to be excited by music that has no human in the loop, nor any connection to the physical world. We are already drowning in music, you can turn on Spotify and have enough music to fill a lifetime. Yet new music is still being produced, why? Because music is ultimately a psychological experience, the human connection is a not-insubstantial part of the experience. There’s a place for AI in music but it has to be white box, there needs to be scope for a human to jump in there, modify things, and make it their own. Otherwise, who will care?
- flangola7 3y agoWe already have infinite streams of Seinfeld
- renewiltord 3y agoFor many genres of music I like, it did a terrific job. I could listen to a few hours of this Chinese instrumental music Electronic remix. And an infinite stream of music is exactly what I want. I don't want to curate or search. I want to feel and ask and get.
- fullshark 3y agoAgreed, reminds me of the MIDI era
- laratied 3y ago[dead]