15 ms·
Stable-Audio-Demo
- lbourdages 3y agoThis is right into the "uncanny valley" of music. It definitely sounded "like music", but none of it is what a human would produce. There's just something off.
- otabdeveloper4 3y agoAI pictures are the same. We are more tolerant of six fingered-pictures with missing limbs, for some reason.
- lbourdages 3y agoWe're used to drawings, 3D renders, etc. There's no such thing as "artificial music" - at the very least, not since electronic music has become mainstream.
- bane 3y agoThe overall audio quality sounds pretty good and it seems to do a good job of sustaining a consistent rhythm and musical concept. But I agree there's something "off" about some of the clips. - The rave music sounds great. But that's because EDM can be quite out there in terms of musical construction. - The guitar sounds weird because it doesn't sound like chords a human hand can make on a tuning nobody tunes their guitar to - with a strange mix of open and closed strings that don't make sense. I think the restrictions of what a guitar can do aren't well understood by the model. - The disco chord progression is bizarre. It doesn't sound bad, but it's unlikely to be something somebody working in the genre would choose. - meditation music - I mean, most of that genre may as well just be some randomized process - drum solo - there's some weird issues in some of the drum sounds, things like cymbals, rides and hats changing tone in the middle of a note, some of the toms sound weird, it sounds like a mix of stick and brush and stick and stick and brush all at the same time...it's sort of the same problem the solo guitar has where it's just not produced within the constraints of what a drum player can actually do on an instrument made of actual drums - sound effects, all are pretty good, a little chunky and low bit-rate or low sample-rate sounding, there's probably something going on in the network that's reducing the rate before it gets build back up. There's a constant sort of reverb in all of the examples I honestly can't say I prefer their model over some of the musicgen output even if their model is doing a better job at following the prompts in some cases. All of the models have a very low bitrate encoding problems and other weird anomalous things. Some of it reminds me of the output from older mp3 encoders, where hihats and such would get very "swishy" sounding. You can hear some of it in the autoencoder reconstructions, especially the trumpet and the last example. However, in any case, I'm actually glad in some ways to see the progress being made in this area. It's really impressive. This was complete science fiction only a very few years ago.
- darkwater 3y ago> - drum solo - there's some weird issues in some of the drum sounds, things like cymbals, rides and hats changing tone in the middle of a note, some of the toms sound weird, it sounds like a mix of stick and brush and stick and stick and brush all at the same time...it's sort of the same problem the solo guitar has where it's just not produced within the constraints of what a drum player can actually do on an instrument made of actual drums And I would say that there is also background noise from time to time, at some point I heard some noise akin to voices. Maybe it is some artifact caused by the training data (many drum solos are performed exclusively live).
- RobinL 3y agoHere is a silly song I generated using suno.ai, which I have found to be incredibly impressive (at least, a small percentage of its outputs are very good, most are bad). I think it's good enough that most humans wouldn't realise it's AI generated. https://app.suno.ai/song/8a64868d-9dd3-46db-91af-f962d4bec8b6 https://app.suno.ai/song/8a64868d-9dd3-46db-91af-f962d4bec8b...
- deleted 3y ago[deleted]
- Agraillo 3y agoVery good for my taste, but I should clarify, I'm obsessed with catchy tunes, as a listener and as a hobby musician, growing my own brainworms from time to time. And I must say that suno.ai is very impressive, in my case semi-ready brainworms are almost always in 30%-50% cases. And what's more important, it's really an inspiration tool for all kinds of tasks, like lyrics polishing or playing-along after track separation. Maybe catchy melodies are not for all, but who can argue with charts when The Beatles, ABBA and Queen were almost always producers of ones.
- comex 3y agoWow. I’m guessing it’s generating MIDI or something rather than synthesizing audio from scratch? Even so, the quality of the score is leaps and bounds better than any of the long-form audio on the Stable Audio demo page (either Stable Audio itself or the other models). The audio model outputs seem to take a sequence of 1 to 3 chords, add a barebones melody on top, and basically loop this over and over. When they deviate from the pattern, it feels unplanned and chaotic and they often just snap back to the pattern without resolving the idea added by the deviation. (Either that or they completely change course and forget what they were doing before.) Yes, EDM in particular often has repetitive chord structures and basic melodies, but it’s not that repetitive. In comparison, from listening to a few suno.ai outputs, they reliably have complex melodies and reasonable chord progressions. They do tend to be repetitive and formulaic, but the repetition comes on a longer time scale and isn’t as boring. And they do sometimes get confused and randomly set off in a new direction, but not as often. Most of the time, the outputs sound like real songs. Which is not something I knew AI could do in 2024.
- dcre 3y agoOne thing I noticed is that when it’s playing chords, it seems a lot more likely than human players to put both major and minor thirds in. This isn’t unheard of — the famous Hendrix chord in “Purple Haze” consists of root, major third, 7th, minor third. But it sounds pretty weird when you do it in every chord.
- ShamelessC 3y agoSo there aren't public weights, is that right? Having trouble finding anything that says one way or the other. edit: Oh okay, didn't realize this was somehow a controversial comment to make. It would have been great if you had answered the question before downvoting but that's fine I suppose.
- NoPedantsThanks 3y ago[flagged]
- grey8 3y agoNope. They did release code for training, inference and fine tuning, but no datasets or weights. See https://github.com/Stability-AI/stable-audio-tools https://github.com/Stability-AI/stable-audio-tools
- ShamelessC 3y agoThanks!
- turnsout 3y agoWonder if it's an IP issue. They don't want every record label coming after them.
- ShamelessC 3y agoYeah that tracks.
- Timwi 3y agoI see what you did there.
- 8n4vidtmkvmk 3y agoThe music is pretty meh but the sound effects are exciting for indie game dev!
- nullandvoid 3y ago<deleted>
- Auracle 3y agoToo bad according to their page you need an enterprise license for even indie games.
- hansonpeter 3y ago[dead]
- romanzubenko 3y agoAs with Stable Diffusion, text prompting will be the least controllable way to get useful output with this model. I can easily imagine midi being used as an input with control net to essentially get a neural synthesizer.
- zone411 3y agoYes. Since working on my AI melodies project (https://www.melodies.ai/ https://www.melodies.ai/) two years ago, I've been saying that producing a high-quality, finalized song from text won't be feasible or even desirable for a while, and it's better to focus on using AI in various aspects of music making that support the artist's process.
- l33tman 3y agoEmad hinted here on HN the last time this was discussed that they were experimenting with exactly that. It will come, by them or by someone else quickly. Text-prompting is just a very coarse tool to quickly get some base to stand on, ControlNet is where the human creativity again enters.
- emadm 3y agoYeah, we build ComfyUI so you can imagine what is coming soon around that. Need to add more stuff to my Soundcloud https://on.soundcloud.com/XrqNb https://on.soundcloud.com/XrqNb
- 3cats-in-a-coat 3y agoText will be an important input channel for texture, sound type, voice type and so on. You can't just use input audio, that defeats the point of generating something new. You can't also only use MIDI, it still needs to know what sits behind those notes, what performance, what instrument. So we need multiple channels.
- numpad0 3y agoIt's crazy that nobody cares. It seems to me that ML hype trends focus on denying skills and disproving creativity by denoising randoms into what are indistinguishable from human generation, and to me this whole chain of negatives don't seem to have proven its worth.
- reissbaker 3y agoThis is incredibly good compared to SOTA music models (MusicGen, MusicLM). It looks like there's also a product page where you can subscribe to use it, similar to Midjourney: https://www.stableaudio.com/ https://www.stableaudio.com/ Sadly it's not open-weight and it doesn't look like there's an API (again like Midjourney): you subscribe monthly to generate audio in their UI, rather than having something developers can integrate or wrap.
- ex3ndr 3y agoThankfully you can train it at home, the bigger question is a data.
- nullandvoid 3y agoI was hoping to use it to generate some sound effects to use in a game I'm working on - but looks like I need an "enterprise license" (https://www.stableaudio.com/pricing https://www.stableaudio.com/pricing) Why does this have a different clause I wonder, and doesn't just fall under "In commercial products below 100,000 MAU"?
- emadm 3y agoDifferent deal with the underlying data holders with revenue share etc
- emadm 3y agoThere is a CC licensed version soon plus API. Models are advancing very fast, will be quite the year for music.
- reissbaker 3y agoAny chance of a commercially licensed version? CC is alright for research but I feel like the real meat of a lot of these models is finetuning. (Or, will the API support finetuning?)
- TillE 3y agoI was briefly excited about the idea of generating sound effects, but those "footsteps" are incredibly bad.
- laborcontract 3y agoI tried generating music on stableaudio.com and, yes, it's bad. However, given the blistering pace of developing in these models, I would not be surprised if these sound incredible in a year or two.
- berkes 3y agoEveryone every time seems to assume a linear (or exponential) curve upwards. But what is the proof for that? I consider it far more likely that we had a breakthrough and now rushing towards the next plateau. Maybe are nearing that. Like in the curve of a PID controller. It's how most or many human improvements go.
- leodriesch 3y agoI'd say most are thinking of Midjourneys success in image generation when talking about this kind of progress.
- berkes 3y agoI'm too. But I still see no evidence that this keeps improving and not plateauing at some (current?) level.
- spacebanana7 3y agoThe plateau we're heading for is getting professional human level output from these models with logarithmic progress. I suspect this is because the underlying production factors like compute, data & model design are steadily improving whilst humans have diminishing sensitivity to output quality. In the game of AI generated photorealistic images or history essays there's not much improvement left to make. Most humans are already convinced by the output of these things.
- andrewstuart 3y agoI felt a great disturbance in the Force, as though all the music licensing lawyers in the USA all cried out at once.
- shon 3y agoPerhaps the disturbance you feel is actually the RIAA moving their Death Star into firing range of Stability.ai
- emadm 3y agostableaudio.com is fully licensed, music is an interesting area https://www.musicbusinessworldwide.com/stability-ai-launches-text-to-music-generator-trained-on-licensed-content-via-a-partnership-with-music-library-audiosparx/ https://www.musicbusinessworldwide.com/stability-ai-launches...
- kouteiheika 3y agoSerious question, I'd genuinely like to know - why? You didn't license the images when training Stable Diffusion, and yet you did for Stable Audio? In both cases the training should either be fair use and legal without any licensing, or be infringing and need licensing. Why is audio different than images? Am I missing something here?
- emadm 3y agoLaw for music is different to other media types
- NoPedantsThanks 3y ago[flagged]
- frizlab 3y agoI know right, what year is this?
- otabdeveloper4 3y agoI do. It's year of the Google (c), like every year. (David Foster Wallace was wrong, there's no way a company of Google's caliber would settle for anything less than a whole decade.)
- shon 3y agoIs it? To me it feels like Google is about where Microsoft was in 2002. Case in point: This thread about using anything other than Chrome…
- XorNot 3y agoWorked fine in Firefox.
- consumer451 3y agoAlso, worked fine in Safari mobile reader mode.
- shon 3y agoWhat’s your preferred browser?
- otabdeveloper4 3y agoAnything that isn't Chrome.
- DaiPlusPlus 3y agoIf I could have my way, NCSA Mosaic.
- andbberger 3y agowake me up when it can write a fugue
- alacritas0 3y agothis can produce some pretty disturbing, but interesting music using the prompt "energetic music, violin, voice, orchestra, piano, minimalism, john adams, nixon in china": https://www.stableaudio.com/1/share/953f079e-d704-4138-904c-5e9ed2d10503 https://www.stableaudio.com/1/share/953f079e-d704-4138-904c-...
- FergusArgyll 3y agoIt reminds me a little of breath of the wild guardian music
- seydor 3y agoFinally, some music from the future
- MrThoughtful 3y agoSo many questions ... They publish the code to train on your own music, but not the weights of their model? So you cannot just upload this thing to some EC2 instance and start creating your own music, correct? Is this the same as https://www.stableaudio.com https://www.stableaudio.com?
- nextworddev 3y agoStabilityAI is just a marketing machine at this point that is praying for an acquisition, since the runway is diminishing
- alacritas0 3y agothis sounds like progress, but it is still very bad except for highly repetitive music like the EDM examples they give, and even then, it still can't get tempo right
- qwertox 3y agoI think we still need the step where the AI learns what a high quality sound library sounds like and then applies the previously learned abilities by triggering sounds of that library via MIDI. That way you'd get perfect audio quality with the creativity of a musical AI.
- eru 3y agoHow would MIDI get you eg a guitar being played dirty? Or some subtle echo that comes from recording in a bathroom?
- arrakeen 3y agothe AI designs and controls the effects chain and mastering too
- qwertox 3y agoIt would use a sampler and for the subtle echo effect add a reverb to the bus. https://www.youtube.com/watch?v=EQdp2QLiSYQ&t=187s https://www.youtube.com/watch?v=EQdp2QLiSYQ&t=187s
- sebzim4500 3y agoYou could have AI do some postprocessing. I think a similaar approach is the future for image generation, you have a model output a 3D scene, use a classical raytracer to do rendering and then have a final model apply corrections to achieve photorealism.
- jchw 3y agoI've always wished for something like that for image generation AI. It'd be much cooler/more interesting to watch AI try to draw/paint pictures with strokes rather than just magically iterate into a fully-rendered image. I dunno what kind of dataset or architecture you could possibly apply to accomplish this, but it would be very interesting.
- Auracle 3y agoI get what you’re saying, but if you watch Stable Diffusion do each step it’s at least kind of similar. If you keep the same seed but change a detail, often the broad “strokes” are completely the same.
- deleted 3y ago[deleted]
- shon 3y agoInterestingly, Ed Newton-Rex, the person hired to build Stable Audio, quit shortly after it was released due to concerns around copyright and the training data being used. He’s since founded https://www.fairlytrained.org/ https://www.fairlytrained.org/ Reference: https://x.com/ednewtonrex https://x.com/ednewtonrex
- az226 3y agoThat’s an interesting take. But quite the odd stance since he joined Stability and the training of Stable Diffusion was well known.
- doctorpangloss 3y agoFor generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to. Even for rightsholders with tens of millions to hundreds of millions of library items like images or audio snippets, the performance of the encoder or similar feature in text-to-X generative models is too poor on the less than billion tokens of text in the large repositories. This includes Adobe's Firefly. It is also a misconception that large amounts of similar data, like the kinds that appear in these libraries, is especially useful. Without a powerful text encoder, the net result is that most text-to-X models create things that look or sound very average. The simplest way to dispel such issues is to publish the architecture of the model. But anyway, even if it were all true, the only reason we are talking about diffusers, and the only reason we are paying attention to this author's work Fairly Trained, is because of someone training on data that was not expressly licensed.
- sillysaurusx 3y agoIf you require licensing fees for training data, you kill open source ML. That’s why it’s important for OpenAI to win the upcoming court cases. If they lose, they’ll survive. But it will be the end of open model releases. To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.
- jpc0 3y ago> Warning: This website may not function properly on Safari. For the best experience, please use Google Chrome Do better
- popalchemist 3y agoHave you ever heard of an MVP?
- prmoustache 3y agoThat would be pertinent if it wasn't just a static web page with just text and some audio files to be played.
- zamadatix 3y agoReading about it, that ironically seems to be the exact problem Safari has. I mean the page "works" in Safari it's just you get these really random delays to the start of some of the sounds with all sorts of web discussion threads saying different ways to mitigate it on different platforms. I don't really fault them for having the goal to publish a paper and go the extra bit to make a friendly but imperfect webpage instead of being website creators who happen to publish papers on the side.
- pmontra 3y agoBy the way, it does work on Firefox Android. No idea of what there is in Safari that's not standard in Chrome and Firefox.
- Aachen 3y ago...and recommend Firefox is what you meant to say right? :)
- lopkeny12ko 3y ago> We append “high-quality, stereo” to our sound effects prompts because it is generally helpful. It's hilarious that we've discovered you can get better outputs from LLMs by simply nicely telling it to generate better results.
- nine_k 3y agoMaybe sometimes you want an old cassette sound, or even older scratched 78 rpm sound, etc. Computers, as usual, do what you asked them to do, not what you meant.
- ttul 3y agoI find it interesting that they are releasing the code and lovely instructions for training, but no model. They are almost begging anonymous folks to hook the data loader up to an Apple Music account and go nuts. Not that I am suggesting anyone do that.
- zamadatix 3y agoSpeculatively it might have been part of an agreement with they were given the licensed stock audio library from AudioSparx to train on they wouldn't redistribute the resulting model.
- jsiepkes 3y ago> Warning: This website may not function properly on Safari. For the best experience, please use Google Chrome. We've come full circle with the 90's and Internet Explorer. Well I guess this time the dominant browser is opensource so that's atleast something... Can someone please create an animated GIF button for Chrome which says: "Best viewed with Google Chrome"?
- Maxion 3y agoChrome isn't open source, chromium is. Best not to confuse the two.
- schleck8 3y agoChrome and Chromium are virtually identical except for Google services, which aren't required to do anything with the browser except for installing Chrome extensions that can alternatively be sideloaded, so this is nitpicking.
- berkes 3y agoIt's essential nitpicking
- urbandw311er 3y agoJumping in to defend parent comment, there’s nothing Open Source about Google Chrome and it’s highly relevant in this context because they are notorious for putting technologies and tracking in there that many people find objectionable.
- forgotusername6 3y agoTangential, but I tried to build chromium the other day but stopped when it said it required access to Google cloud platform to actually build it. If something requires a proprietary build system, does it matter that it's open source?
- nolist_policy 3y ago
- ecmascript 3y agoJust a few days ago I was down voted for stating AI will be better in creating music than human would be: https://news.ycombinator.com/item?id=39273380#39273532 https://news.ycombinator.com/item?id=39273380#39273532 Now this is released and now I feel I got grist to my mill. Sure it still kind of sucks, but it's very impressive for a _demo_. Remember that this tech is very much in it's infancy and it's very impressive already.
- larschdk 3y agoI don't find this music to be good in any way. It sounds interesting over a few notes, but then completely fails to find any kind of progression that goes anywhere interesting, never iterating on the theme, never teasing you with subtle or surprising variation over a core theme, no built-ups or clear resolution. Very annoying to actually listen to.
- webprofusion 3y agoMusic is perfect for AI generation using trained models, because artists have been copying each other for at least the past 100 years and having a computer do it for you is only notionally different. Sure a computer can never truly know your pain, but it can copy someone else's.
- deleted 3y ago[deleted]
- kleiba 3y agoMy son suggested to play "Calm meditation music to play in a spa lobby" and "Drum solo" at the same time - sounds pretty good, actually...
- Jeff_Brown 3y agoThat's some pretty advanced musicality.
- PeterStuer 3y ago"Gen AI is the only mass-adoption technology that claims it's Ok to exploit everyone's work without permission, payment, or bringing them any other benefit." Is it? What about the printing press, photography, the copier, the scanner ... Sure, if a commercial image is used in a commercial setting, there is a potential legal case that could argue about infringement. This should NOT depend on the production means, but on the merit of the comparisons of the produced images. Xerox should not be sued because you can use a copier to copy a book (trust me kids, book copying used to be very, very big). Art by its social nature is always derivative, I can use diffusion models to create uncontestably original imagery. I can also try to get them to generate something close to an image in the training set if the model was large enough compared to the training set or the work just realy formulaic. However. It would be far easier and more efficient to just Google the image in the first place and patch it up with some Photoshop if that was my goal.
- wnkrshm 3y agoBut the social nature of art also means that humans give the originator and their influences credit - of course not the entire chain but at least the nearest neighbours of influence. While a user of a diffusion generator does not even know the influences unless specifically asked for. Shoulders of giants as a service.
- webmaven 3y ago> Xerox should not be sued because you can use a copier to copy a book (trust me kids, book copying used to be very, very big). The appropriate analogy here isn't suing Xerox, but suing Kinko's (now FedEx Office). And it isn't just books, but other sorts of copyrighted material as well, such as photographs, which are still an issue.
- haswell 3y ago> Art by its social nature is always derivative, I can use diffusion models to create uncontestably original imagery How are you defining “uncontestably original” here? The output could not exist if not for the training set used to train the model. While the process of deriving the end result is different than the one humans use when creating artwork, the end result is still derived from other works, and the degree of originality is a difference of degree, not of kind when compared to human output. (I acknowledge that the AI tool is enabled by a different process than the one humans use, but I’m not sure that a change in process changes the derivative nature of all subsequent output). As a thought experiment, imagine that assuming we survive, after another million years of human evolution, our brains can process imagery at the scale of generative AI models, and can produce derivative output taking into account more influences than any human could even begin to approach with our 2024 brains. Is the output no longer derivative? Now consider the future human’s interpretation of the work vs. the 2024 human’s interpretation of the work. “I’ve never seen anything like this”, says the 2024 human. “The influences from 5 billion artists over time are clear in this piece” says the future human. The fundamental question is: on what basis is the output of an AI model original? What are the criterion for originality?
- slicerdicer1 3y agoobviously someone shadowy and non-corporate (eg. an artist) just needs to come out and make a model which includes promptable artist/producer/singer/instrumentalist/song metadata. describing music without referring to musicians is so clunky because music is never labelled well. of course saying "disco house with funk bass and soulful vocals, uplifting" is going to be bland. Saying "disco house with nile rodgers rhythm guitar, michael mcdonald singing, and a bassline in the style of patrick alavi's power" is going to get you some magic
- ever1337 3y agoso this model can only ever understand music which is classified, described, labelled, standardized. and recombine those. sounds boring, sounds like the opposite of what (I would like to believe) people listen to music for, outside of a corporate stock audio context.
- gregorvand 3y agoNot trying to knock the progress here, impressive. As a drummer, 'drum solo' is about as boring as it gets and some weird interspersing sounds. So, it depends on the intended audience. FWIW the sound effects also are not 'realistic' to my ear, at the moment. But again, the progress is huge, well done!
- toxik 3y agoYeah the drum solo really highlights how badly the model missed the point in a drum solo. I'm not a drummer, but this is just not pleasing to hear. Sounds like somebody randomly banging drums more or less in tempo. It does okay with muzak-type things though, which I guess tracks with my expectations.
- ZoomZoomZoom 3y agoAs a drummer, the 'drum solo` was surprisingly interesting to listen to, if you consider it happening over a stable 4/4 pulse. The random-but-not-quite nature of the part makes for very unconventional rhythmic patterns. I'd like to be able to syncopate like this on the spot. Don't ask me to transcribe it. Tempo consistency is great. Extraneous noises and random cymbal tails show the deficiency of the model though.
- redman25 3y agoI think I was more disappointed by the music samples not having any transitions. Most songs have key changes and percussion turnovers.
- pier25 3y agoI agree. It's an impressive effort but it's still very far from being able to generate viable music/sound. There are already millions of library music tracks and sound effects available which sound a lot better. It's going to take a huge investment in gen AI to compete with that and I don't think it makes economic sense (unlike text or images).
- TrackerFF 3y agoNow, if they can also generate MIDI-tracks to accompany - that'd be great. That would add some much-needed levels of customization.
- seydor 3y agoTrying to describe music with words is awkward! We need a model that is trained on dance
- Jeff_Brown 3y agoOr architecture.
- zdimension 3y agoThe few examples I was able to play are very promising, unfortunately the host seems to be getting some sort of HN-hug, because all the audio files are buffering every other second -- they seem to throttle at 32 KiB/s.
- Jeff_Brown 3y agoMusic without changes is boring. I enjoyed the much less stable results of OpenAI's JuleBox (2021?) more than any music AI to come since. Their sound quality is better but they only seem to produce one monotonous texture at a time.
- coldcode 3y agoAs a musician, I found the pieces unremarkable. Of course, a lot of contemporary music is forgettable as well, as people try to create songs that all sound like hits but, in doing so, create uninteresting songs. I wonder what music the model is based on. I suppose for game music/sounds, perhaps its good enough?
- 3cats-in-a-coat 3y agoThe reconstruction demo is in effect an audio compression codec. And I bet it makes existing audio codecs look like absolute toys.
- deleted 3y ago[deleted]
- emadm 3y agoThis is part of a paper on the prior version of the model: https://x.com/stableaudio/status/1755558334797685089?s=20 https://x.com/stableaudio/status/1755558334797685089?s=20 https://arxiv.org/abs/2402.04825 https://arxiv.org/abs/2402.04825 Which outperforms similar music models. The pace is accelerating and even better ones are coming with far greater cohesion and... stuff. Will be quite the year for music.
- emadm 3y agoParticularly interesting with the scaled up version of https://www.text-description-to-speech.com https://www.text-description-to-speech.com Do try https://www.stableaudio.com https://www.stableaudio.com for rights licensed model you can use commercially.
- nprateem 3y agoThe problem with music generation is difficulty in editing. Photos and text can be easily edited, but music can't be. Either the piece needs to be MIDI, with relevant parameterisation of instruments, or a UI creating that allows segments of the audio to be reworked like in-painting.
- Jeff_Brown 3y agoWhat's the easiest way you found for using AI to edit photos? I was just yesterday looking at the openai dolly 3 API and it feels pretty limited. For instance, in the picture I have of a fisherman with too many fishing lines hanging down from his fishing rod, I'd like to just point it at the extra fishing lines and say make these go away, but there's no way to do that.
- Wistar 3y agoA small point: Needs to be in something other than 44.1kHz. The two to which they make comparisons are at either 32kHz or 48kHz, both of which are friendlier for video work, something for which I think AI audio will be used a lot.
- m3kw9 3y agoLots of work left to do man
- williamcotton 3y agoNone of these tools are even remotely useful to me unless I can give grooves, chord changes and melodic themes. It’s just a glorified loop library at this point!
- joecarrot 3y agoWhy are AI developers so goddamned keen on having it make art, one of the few kinds of work that human beings actually LIKE doing? We could use AI to be a CPA, or to write citations for a paper, but noooo, AI has to be a painter and a musician. It's almost like the software developers are jealous that someone out there is having a good time and want to take it from them. Also miss me with that 'AI enables me (a scrub) to make art I couldn't otherwise because I don't want to learn how to do it'. You are lazy. Congrats on finding a high horse about your laziness.
- notfed 3y agoI assume you're just being tongue-in-cheekfully dramatic, but the answer of course, is that there are AIs for those things, but they're under much less demand and are much less controversial.
- eutropia 3y agoI think the development of Generative models for images and audio has more to do with the fact that Computer Vision research goes back decades, and the same systems that originally recognized and labeled images or audio were tweaked to invert the process - and it became naturally an intriguing topic of development precisely because creation is seen as an innately human thing. Beyond that, I'd speculate that the reason we keep seeing developments in "the arts" (though I disagree that an AI can make art, even if it can make beautiful images or music) is because there's no readily-agreed-upon value for that task. An AI CPA has a specific economic value, but is also a commodity service that no one wants unless they need it. Since there's a clearly comparable cost for needed CPA services, then naturally creating an AI system to do it has a readily comparable market price. People aren't going to make that AI system unless they can do it in way that will make be an improvement as compared to that existing service and price. I think "just because" has always been a justifiable reason for humans creating beauty (not the same as making art), so it works for research projects better than building a better mousetrap.
- joecarrot 3y agoThanks for the thoughtful reply! You've given me some stuff to think about
- exword76nick 3y agorock band guitar lead solo performance