3 ms·
I'm trying to make a neural audio codec using a variety of misguided methods. One I am using ESNs wrong spreading leak rates in a logarithmic fashion acting lik
by robviren 10mo ago
I'm trying to make a neural audio codec using a variety of misguided methods. One I am using ESNs wrong spreading leak rates in a logarithmic fashion acting like a digital cochlea. The other is trying to do the same with a complex mass-spring-damper system to simulate the various hairs of the cochlea as well. Both approaches make super interesting visuals and appear to cluster reasonably well, but I am still learning about RVQ and audio loss (involves GANs and spectral loss). I kinda wanna beat SNAC if I can.
- Moosdijk 10mo agoDo you have a log available somewhere?
- iFire 10mo agoReminds me of https://github.com/RobViren/kvoicewalk https://github.com/RobViren/kvoicewalk where people take voice clips and train a text to speech using random walks. Not related, misguided methods :D
- Moosdijk 10mo agoWell, it’s the same author so it is kind of related.
- robviren 10mo agoI keep everything in my self hosted gitea. Just made it public. https://gitter.swolereport.com/robviren/cspace https://gitter.swolereport.com/robviren/cspace
- Moosdijk 10mo agoThanks, I’ll check it out Edit: timed out