6 ms·
Nvidia Fugatto: "World's Most Flexible Sound Machine"
- popalchemist 2y agoWill they be releasing weights?
- SonOfLilit 2y agoThe description is amazing, but the demo video feels underwhelming. Available music generation models sound much more musical and have much better diction on vocals.
- codedokode 2y agoThis might be due to quality of the dataset because Nvidia seems to be not using copyrighted commercial recordings (if I read their paper properly). It is difficult to compete with those who have used larger and higher quality dataset without permission.
- SushiHippie 2y agoAnyone knows what melody this is at 2:07 in the Video? https://youtu.be/qj1Sp8He6e4?t=2m7s https://youtu.be/qj1Sp8He6e4?t=2m7s
- pil0u 2y agoAm I the only one feeling weird about the image they chose to illustrate the article? I'm not a professional in that field but I would probably feel offended if coding assistants were presented with a monkey in front of a computer.
- gloflo 2y agoIt's low quality AI slob. That's already offensively disrespectful to the readers on its own.
- throw310822 2y agoUh? Apart from the fact that the symbolism between a monkey and a cat is entirely different, I imagined it was because gato/ gatto means cat in Spanish/ Italian.
- Klaster_1 2y agoSame in greek! I was pleasantly surprised to see a cat, I think this was a nice touch.
- ipsum2 2y ago[flagged]
- Cumpiler69 2y ago>I would probably feel offended if coding assistants were presented with a monkey in front of a computer As a professional monkey in front of a computer, I feel offended.
- gus_massa 2y agoNah. Cats are cool. S/he has a smart cool look. Most people would like it. (Not everyone.) Disclainmer: I prefer dogs, they are more friendly and even part of the family, but I have to recognize that cats look more cool.
- dagw 2y agoCalling a musician a "cool cat" has been a slang term of high praise for jazz musicians since at least the 50s.
- camillomiller 2y agoAnother day, another model made by engineers who think their technical prowess needs absolutely no understanding of the subtlety of human creativity
- ahofmann 2y agoWhile this might be a technical breakthrough, none of the examples sounded any good. Every aspect of the provided sounds are bad. The music sounds muffled and badly mixed. The generated beat isn't a beat that grooves, or has anything interesting in it. The barking saxophone sounded just bad. The voices sounded somewhat convincing. In general I think that with ai generated audio it is much more noticeable how utterly bad everything is, that ai generates. I already absolutely hate the two AI voices that are in a lot of YouTube videos and are a reason for me to close the Video immediately most of the time.
- RobinL 2y agoWith apologies for the X link, here is an example from Suno which felt very musical to me: https://x.com/sunomusic/status/1857501332560818342 https://x.com/sunomusic/status/1857501332560818342 Here's another example on the Suno website: https://suno.com/song/fc991b95-e4e9-4c8f-87e8-e5e4560755e7 https://suno.com/song/fc991b95-e4e9-4c8f-87e8-e5e4560755e7
- numpad0 2y agoI don't find any problems whatsoever in those audio, but I'm not an avid music listener, so out of intuition I'm making a guess that there's same underlying issue as image generation happening: AI makes technically horrible and rage-inducing fillers that lack high level semantic structure, but average people has no words nor experience to assess and describe what's going on.
- ahofmann 2y ago> I don't find any problems whatsoever in those audio I think this is why there is no real, powerful protest against all that generated stuff. Only the people, that care, are able to articulate what's wrong with it. To me, all of AI generated content sounds horrible. To almost everyone else, this sounds ok. So we will see and hear more of this generated stuff. We are in the middle of the enshittification of all consumable media.
- 2y ago
- olau 2y agoI would love to see a model focusing on making virtual instruments. There are sample-based virtual instruments, but they do miss some subtleties, and there are physics-based ones where some subtleties are preserved, but generally worse sounding because actually modelling real hardware evolved over centuries is really difficult. Even hardware-based instruments like electric guitars/violins/cellos etc. generally sound distinct from and less interesting than their acoustic counterparts. Electric guitar players seem to use various amplifier tricks to make up for that, and that's now a big separate instrument. But I think the point stands.
- ulbu 2y agoI concur. So much focus on grandiose ideas when there are so many low-hanging fruit around.
- insomagent 2y agohttps://www.neuralampmodeler.com/ https://www.neuralampmodeler.com/
- ZoomZoomZoom 2y agoMost of audio and music AI have wrong incentives and are moving in a different direction professionals need. Almost all publicized innovations in the sphere are complex one-stop-shop solutions which aim to completely replace as many members of the creative process as possible. It's a corporate dream: a thing that spews barely-passable, generic mush that's totally aligned with demands of the decision makers, but has zero opinion, zero ambition, zero professional pride and no need to uphold its ethical and aesthetic standards and its own reputation whatsoever. Instead of tools for the creatives we have systems that generate complete tracks from tinder chat logs. On the other hand, there's still no publicly available audio style transfer with even remotely usable quality (that thing from Google is abysmal). All I want for starters is something that turns a slightly distorted, over-reverberated and not-perfectly intonated flute recording that a client sends me into a clean workable track. I don't even ask for it to turn it into koto or marimba or whatever you think is a cool demonstration case! Sorry for the rant, but it's all very frustrating and alarming.
- tgv 2y agoBarely passable, indeed. And then to imagine the MBAs are indeed going to fire staff and downsize contractors because of this. More money for them: is that the incentive here?
- ecocentrik 2y agoIs it being trained on noticeably compressed audio or is it just outputting highly compressed audio? Can someone explain what the benefits of either would be outside of specifically asking for the sound of audio compression artifacts? Like others have pointed out, existing generative music services already output much higher fidelity audio.
- varispeed 2y agoPlease don't put headphones over a cat's head and especially don't play any loud music!
- codedokode 2y agoIf you use it for work, AI might be ok, but generating a guitar or piano track is zero fun compared to playing a real instrument (even if AI track sounds better). I think we should not forget this part too. But what about an AI guitar that automatically frets the strings properly if you don't press them hard enough? Or an AI piano which shifts the keyboard when it sees that you are about to hit the wrong key? Many instruments require lot of practice before you can produce acceptable sound. Can AI help with this?
- norir 2y agoNot only do the instruments require practice to sound good (I've been playing electric bass for three years and am just beginning to sound better than bad), but a huge part of the process is learning to listen to the instrument and make adjustments. The beauty is that you can immediately hear the result of the adjustment. If it sounds better, you keep it. Otherwise you move until you get closer to what you're looking for. With a prompt based ai tool, it is not possible to make low latency adjustments. Even if you could, how would you articulate the subtle adjustment to the llm? My sense is that contrary to marketing, ai tools will be most useful to people who already have musical skill and will actively subvert musical development in most people who rely on it too early in their process.
- olup 2y agoThey say the models are under 3b parameters. If only for voice generation it sounds pretty good, no ?