7 ms·
All the AI music I’ve heard so far has a really unpleasant resonant quality to it. Why is that? Can it be removed?
by newobj 4y ago
All the AI music I’ve heard so far has a really unpleasant resonant quality to it. Why is that? Can it be removed?
- hyperbovine 4y agoPresumably for similar reasons that the vast majority of AI generated art and text is off-puttingly hideous or bland. For every stunning example that gets passed around the internet, thousands of others sucked. Generating art that is aesthetically pleasing to humans seems like the Mt. Everest of AI challenges to me.
- andybak 4y agoI think your comment is off-topic to the post you are replyng to. That wasn't asking about the general aesthetic quality - more about a specific audio artifact. > For every stunning example that gets passed around the internet, thousands of others sucked. From personal experience this is simply untrue. I don't want to debate it because you seem to have strong feelings about the topic.
- hyperbovine 4y agoEven if you remove the artifact, the exact same comment applies. It generates a somewhat less interesting version of elevator music. This is not to crap on what they did. As I said, they underlying problem is extremely difficult and nobody has managed to solve it. I don't feel strongly about this topic at all.
- indigochill 4y ago> It generates a somewhat less interesting version of elevator music. This iteration does, but that's an artifact of how it's being generated: small spectograms that mutate without emotional direction (by which I mean we expect things like chord changes and intervals in melodies that we associate with emotional expressions - elevator music also stays in the neutral zone by design). I expect with some further work, someone could add a layer on top of this that could translate emotional expressions into harmonic and melodic direction for the spectrogram generator. But maybe that would also require more training to get the spectrogram generator to reliably produce results that followed those directions?
- jameshart 4y agoThe vast majority of human generated art is hideous or bland. Artists throw away bad ideas or sketches that didn’t work all the time. Plus you should see most of the stuff that gets pasted up on the walls at an average middle School.
- ROTMetro 4y agoHard disagree. The average middle school picture will have certain aspects exaggerated giving you insights into the minds eye of the creator, how they see the world, what details they focus on. There is no such minds eye behind AI art so it's incredibly boring and mundane, no matter how good a filter you apply on top of it's fundamental lack of soul or anything interesting to observe in the picture beyond surface level. It's great for making art for assets for businesses to use, it's almost a perfect match, as they are looking to have no controversial soul to the assets they use, but lots of pretty bubblegum polish.
- jameshart 4y agoAnd the vast majority of professionally produced artwork is for business use. It’s packaging design or illustration or corporate graphics or logos or whatever. I don’t get the objection.
- antipotoad 4y agoPerhaps most of the AI art out there (that honestly represents itself as such) is boring and mundane, but after many hours exploring latent space, I assure you that diffusion models can be wielded with creativity and vision. Prompting is an art and a science in its own right, not to speak of all the ways these tools can be strung together. In any case, everything is a remix.
- dwringer 4y agoI have to agree, the act of coming up with a prompt is one and the same with providing "insights into the minds eye of the creator, how they see the world, what details they focus on" - two people will describe the same scene with completely different prompts.
- adamsmith143 4y agoNot sure about this. Models like Midjourney seem to put out very consistently good images.
- blueboo 4y ago> For every stunning example that gets passed around the internet, thousands of others sucked …implying there may be an art to AI art. Hmm. Meanwhile, the degree to which it is off-puttingly hideous in general can be seen in the popularity of Midjourney — which is to observe millions of folks (of perhaps dubious aesthetic taste) find the results quite pleasing.
- woah 4y agoYou're probably talking about the artifacts of converting a low resolution spectrogram to audio.
- wdfx 4y agoCan the spectrogram image be AI upscaled before transforming back to the time domain?
- malka 4y agoYes it exists: https://ccrma.stanford.edu/~juhan/super_spec.html https://ccrma.stanford.edu/~juhan/super_spec.html But the issue is not that the spectrogram is low quality. The issue is that the spectrogram only contains the amplitude information. You also need phase information for generating audio from the spectogram
- mcbuilder 4y agoInteresting, can't you quantize and snap to a phase that makes sense to create the most musical resonance?
- waltbosz 4y agoWhat happens if you run one of the spectrogram pictures through an upscaler for images like ESRGAN ?
- syntheweave 4y agoIt sounds kind of like the visual artifacts that are generated by resampling in two dimensions. Since the whole model is based on compressing image content, whatever it's doing DSP-wise is more-or-less "baked in", and a probable fix would lie in doing it in a less hacky way.
- recursive 4y agoThe link is down now, so I don't know about this one. But most generated music is generated in the note domain, rather than the audio domain. Any unpleasant resonance would introduced in the audio synthesis step. And audio synthesis from note data is a very solved problem for any kind of timbre you can conceive of, and some you can't.
- crubier 4y agoI think this is because the generation is done in the frequency domain. Phase retrieval is based on heuristics and not perfect, so it leads to this "compressed audio" feel. I think it should be improvable
- antognini 4y agoI've done some work on AI audio synthesis and the artifacts you're hearing in these clips are coming from the algorithm that is used to go from the synthesized spectrogram to the audio (the Griffin-Lim algorithm). Audio spectrograms have two components: the magnitude and the phase. Most of the information and structure is in the magnitude spectrogram so neural nets generally only synthesize that. If you were to look at a phase spectrogram it looks completely random and neural nets have a very, very difficult time learning how to generate good phases. When you go from a spectrogram to audio you need both the magnitudes and phases, but if the neural net only generates the magnitudes you have a problem. This is where the Griffin-Lim algorithm comes in. It tries to find a set of phases that works with the magnitudes so that you can generate the audio. It generally works pretty well, but tends to produce that sort of resonant artifact that you're noticing, especially when the magnitude spectrogram is synthesized (and therefore doesn't necessarily have a consistent set of phases). There are other ways of using neural nets to synthesize the audio directly (Wavenet being the earliest big success), but they tend to be much more expensive than Griffin-Lim. Raw audio data is hard for neural nets to work with because the context size is so large.
- xnzakg 4y agoConsidering Stable Diffusion generates 3-channel (RGB) images, maybe it would be possible to train it on amplitude and phase data as two different channels?
- antognini 4y agoPeople have tried that, but the model essentially learns to discard the phase channel because it is too hard for it to learn any useful information from it.
- deleted 4y ago[deleted]
- techdragon 4y agoGot any citations... that sounds like a fascinating thing to read about.
- gdubs 4y agoThe first ever recordings had people shouting to get anything to register. They sounded like tin. Fast forward to today. Looking back at image generation just a year or two ago and people would have said similar things. Not hard to imagine the trajectory of synthesized audio taking a similar path.