7 ms·
The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic
by iandanforth 3y ago
The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable.
While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to output full images, but either layers or brush strokes and tool pallet usage. Same for organizations that have all the component tracks of music to work with.
- Jeff_Brown 3y agoThat raises an interesting difference between cleaning AI-generated sound and cleaning ordinary recordings. In an ordinary recording, there is an objective reality to discover -- a certain collection of voices was summed to create a signal. With (most? the best?) existing AI audio generation, the waveform is created from whole cloth, and extracting voices from it is an act of creation, not just discovery. I've come across AI-generated music that outputs something like MIDI and controls synthesizers. Its audio quality was crystal-clear, but the music was boring. That's not to say the approach is a dead-end, of course -- and indeed, as a musician, the idea of that kind of output is exciting. But getting good data to train something that outputs separate MIDI-ish voices seems much harder than getting raw audio signals.
- fnordpiglet 3y agoGenerative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.
- jskherman 3y agoI believe Spotify's Basic Pitch[0] is already some work towards building something like this. [0]: https://basicpitch.spotify.com/about https://basicpitch.spotify.com/about
- MrCheeze 3y agoIt has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971 https://nitter.net/jwthickstun/status/1669726326956371971
- radarsat1 3y ago> Generative models can certainly create midi, but no one has done it yet. Note sequence generation from statistical models has a long history, at least as long if not longer than text generation. Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation. And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and transformers. [0]: https://www.researchgate.net/publication/345915209_A_Comprehensive_Survey_on_Deep_Music_Generation_Multi-level_Representations_Algorithms_Evaluations_and_Future_Directions https://www.researchgate.net/publication/345915209_A_Compreh...
- fnordpiglet 3y agoYes, in fact I think at some point everyone has written their own Markov generators or at least run dissociative press. But we’ve really only seen meaningfully high quality output over the last few years.
- radarsat1 3y agoI think it depends on how you define that. People were quite happy with HMM-based MIDI generators that could generate Beethoven- or Mozart-like sequences 10, maybe even 15 or 20 years ago. But of course other people pointed out the problems of it being boring eventually. Then LSTMs improved long-term dependencies and people were impressed by the improved quality of generating whole musical pieces. But still others thought it was not good enough. Then the goalposts moved again with transformers and neural vocoders and now we want top-40 direct audio generation. And these latest systems can kind of sort of do it! But still there are people who demand better. And so on, things will continue to improve. Progress only moves as fast as expectations, and expectations move with technology. Music is not special in this respect. So you could say at any given time in the past that some people "see meaningfully high quality" and others are disappointed. You see exactly both these sides of the spectrum even now with text-to-image and text-to-audio technology.
- fassssst 3y agoDo you know if anyone has tried training a text-to-music or text-to-midi model where the training data includes things like emotion labels for each note interval or chord progression?
- TheActualWalko 3y agoWe’ve done it! wavtool.com
- fnordpiglet 3y agoThat’s really neat. How long have you been working on this?
- TheActualWalko 3y agoThanks! It grew out of an old side project. Been full time on it since December.
- dylan604 3y ago>For example, if I were Adobe I would be training models, not to output full images, but either layers or brush strokes and tool pallet usage. Same for organizations that have all the component tracks of music to work with. I really like this idea. Creating new tools for artists to use to create rather than whatever we're accepting as use now. The use of current full image creation is boring to me in the same way the choice of invisibility as a super power is. The invisibility is ultimately going to slide into pervy tendencies, just like deep fakes will slide in the same way or some other inappropriate use.
- fnordpiglet 3y agoThere are a lot of Lora models that are being made to generate textures, maps, diagrams, backgrounds, etc. You don’t need to wait for adobe, open source models like stable diffusion let you do whatever you think is useful. I’d look to the open source world for creative innovation. Adobe is just doing what’s on the product management roadmap.
- gabereiser 3y agoYessssss! I thought about MusicGAN and Markov chains last night thinking “Why can’t we just codify all chords and use a GAN to generate markov chains on chords of a key and have AI generate instruments and waveform from those chains?” IANA researcher but in my head, that sounded logical.
- TylerE 3y agoThat's existed for decades. It's called Band in a Box. It's also cheezy as hell.
- gabereiser 3y agolol, no. Not autogenerate midi (although their latest versions of BiaB are pretty darn good now) but generate waveforms together. It would be similar to having AI generate whole scores of music but ensuring it's all in sync and in key. Not taking sample database of 88 sound files and triggering them when the midi-note strikes.
- TylerE 3y agoThat's not how BiaB works at all. It has all kinds of patterns built into it. So, it knows, how to generate, say, a bluegrass bassline in a given key. There are plenty of ways to play back MIDI with high sound quality, including feeding it into an AI-driven VST like NotePerformer.
- gabereiser 3y agoThen explain why you categorize it as cheesy? Sounds like it’s pretty cool.
- 93po 3y agohe's saying the output is cheesy. it sounds like stuff you'd hear on a demo track for a kid's toy piano
- miohtama 3y agoHaving music editable for human post production is necessary for most professional adoption. Generating MIDIs would make much more sense than generating raw audio. This is what we do with AI images: you can fix them in Photoshop, etc. You cannot do this for raw audio due to how music is produced.
- waffletower 3y agoBuild or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.
- iainctduncan 3y agoMidi is not lossy, midi is symbolic. There's a huge difference.
- Applejinx 3y agoNo, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.
- gamblor956 3y agoMIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?
- waffletower 3y agoThe problems have long been known and articulated: http://www.music.mcgill.ca/~gary/courses/papers/Moore-Dysfunctions-CMJ-1988.pdf http://www.music.mcgill.ca/~gary/courses/papers/Moore-Dysfun...
- waffletower 3y agoHopefully, the entire industry will NOT move in such a schematic and lossy direction. Use separate tools to analyze audio streams please. Don't throw the timbre baby out with the bathwater. MusicGen utilizes a tokenized transformer model for music, which is attractive for symbolic translation use cases. However, the overall audio quality is far more lossy than the examples you hear from Stable Audio. I believe that symbolic representation should not be a foundational approach to adequately represent and generate rich audio signals.
- tech_ken 3y agoI was wondering the same thing, definitely seems like generating the raw waveform runs into all kinds of weird issues (like they touched on in this post). I would imagine that training data would be a serious chokepoint here. Given how much discourse is currently kicking off around the intellectual property rights of just the final product (the mastered track), I can't imagine many musicians would be eager to share what is effectively the "proof of ownership" (track stems or MIDIs).
- schazers 3y agoI strongly agree about generating "editables" rather than finalized media. In fact, that's why text generators are more useful than current media generators: text is editable by default. Here's a tweetstorm about it: https://x.com/jsonriggs/status/1694490308220964999?s=20 https://x.com/jsonriggs/status/1694490308220964999?s=20
- waffletower 3y agoAudio is definitely editable. While generative audio is new I am hopeful that a host of interesting applications will emerge (audio2audio etc.) within its ecosystem. Promising signal separation (audio to STEMs) and pitch detection tools already exist for raw audio signals. If you want to force Stability to focus on symbolic representations (such as severely lossy MIDI) I hope you can instead first try adapting to tools that work fundamentally with rich audio signals. Perhaps there will be room for symbolic music AI and perhaps Stability will even develop additional models that generate schematic music, but please please don't sacrifice audio generality for piano roll thinking alone. LORAs will undoubtedly be usable to generate more schematic audio via the Stable Audio model -- I imagine they could be easily purposedly to develop sample libraries compatible with DAW (digital audio workstation), sequencer and tracker production workflows.
- visarga 3y agoTrain the model with midi notes as text in the prompt and the audio as target. It will learn to interpret notes.
- waffletower 3y agoNot all music is well represented with notes, nor are audio datasets with high-quality note representations readily available. But I guess if you work hard enough you can get close: https://www.youtube.com/watch?v=o5aeuhad3OM https://www.youtube.com/watch?v=o5aeuhad3OM My example still sounds like the chiptune simulation that it is, however.
- 3y ago
- rhelsing 3y agoI am approaching this from the symbolic angle via MIDI at neptunely (https://neptunely.com/ https://neptunely.com/)