8 ms·
This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you p
by valdiorn 4y ago
This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output.
Absolutely blows my mind.
- 323 4y agoYou can also add another neural-network to "smooth" the spectrogram, increase the resolution and remove artefacts, just like they do for image generation.
- bckr 4y agoPretty sure that's how RAVE works
- hyperbovine 4y agoWasn't this Fraunhofer's big insight that led to the development of MP3? Human perception actually is pretty forgiving of perturbations in the Fourier domain.
- ComplexSystems 4y agoIn very limited situations. You can move a frequency around (or drop it entirely) if it's being masked by a nearby loud frequency. Otherwise, you would be amazed at the sensitivity of pitch perception.
- dehrmann 4y agoThe easy example of this is playing a slightly out of tune guitar, or a mandolin where the strings in the course aren't matched in pitch perfectly. You can hear it, and it's just a few cents off.
- w-m 4y agoYou probably mean Karlheinz Brandenburg, the developer of MP3, who worked on psychoacoustics. Not completely off though, as he did the research at a Fraunhofer research institute, which takes its name from Joseph von Fraunhofer, the inventor of the spectroscope.
- th0ma5 4y agoDoes the institute not also claim that work?
- w-m 4y agoFair enough. But for me, when talking about `having an insight`, I don't imagine a non-human entity doing that. And to be pedantic (talking about Germans doing research, I hope everyone would expect me to be), the institute is called Fraunhofer IIS. `Fraunhofer` would colloquially refer to the society, which is an organization with 76 institutes total. Although, of course, the society will also claim the work...
- WanderPanda 4y agoBringing the right people together and having the right environment that gives rise to „having an insight“ can be a big part as well.
- humanistbot 4y agoIt's an interesting question, one I hadn't thought of before. But in common language, it sometimes makes sense to credit the institution, others just the individuals. I think may be more based around how much the institution collectively presents itself as the author and speaks on behalf of the project versus the individuals involved. Here is my own general intuition for a few contrasting cases: Random forests: Ko and Breiman's, not really Bell Labs and UC-Berkeley Transistors: Bardeen, Brattain, and Shockley, not really Bell Labs (thank the Nobel Prize for that) UNIX: Primarily Bell Labs, but also Ken Thompson and Dennis Richie (this is a hard one) GPT-n: OpenAI, not really any individual, and I can't seem to even recall any named individual from memory
- seth_ 4y agoAuthor here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.
- Dracophoenix 4y agoAny chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.
- utopcell 4y agoThis already exists [1]. [1] https://www.respeecher.com/ https://www.respeecher.com/
- dannyw 4y agoAre there any open source models with good quality? I had a look around several months ago, and it seems like everything is locked behind SaaS APIs.
- echelon 4y agoJames Earl Jones: https://fakeyou.com/tts/result/TR:9ek4x6eb80kq49e94grnhctk4gmcg https://fakeyou.com/tts/result/TR:9ek4x6eb80kq49e94grnhctk4g... Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnpzwf https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp... Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!
- pardon_me 4y agoExcellent work! Singing would be amazing - karaoke can finally sound good :p Have you released a tool for volumetric capture? I'm applying this to LED lighting fixture setup for tv/film/live shows and 3D positioning is the last step to fully automated configuration. My goal is real-time sync between 3D model and real world.
- TaupeRanger 4y agoIt's...not effective though. Am I listening to the wrong thing here? Everything I hear from the web app is jumbled nonsense.
- jefftk 4y agoI think the progression from church bells to electronic beats is especially good: https://www.riffusion.com/about/church_bells_to_electronic_beats.mp3 https://www.riffusion.com/about/church_bells_to_electronic_b...
- itronitron 4y agoI think we're at the point, with these AI generative model thingies, where the practitioners are mesmerized by the mechatronic aspect like a clock maker who wants to recreate the world with gears, so they make a mechanized puppet or diorama and revel in their ingenuity.
- andybak 4y agoAnd that's a bad thing? How do you think human endeavours progress other than by small steps?
- gdubs 4y agoLook at GAN art from a few years ago, compared to MidJourney v4.
- bowsamic 4y agoReally? They sound quite clearly like the prompt to me if I “squint my ears” a little