4 ms·
Very cool, and easy to use! Can you give some more info on how you generated the models? I'm also interested in the tech stack you're using to implement this w
by sunsetMurk 6y ago
Very cool, and easy to use!
Can you give some more info on how you generated the models? I'm also interested in the tech stack you're using to implement this webapp... Would love some details!
..What's next?
- lowdose 6y ago> What's next? Text To Video webapp that renders text to video + voice synchronised of famous people. Who wouldn't like to laugh 5X more when social scrolling? The first platform that enables creators with the ability to produce deep fakes of celebs from text that they can broadcast as HQ video content to their audience will kill both Youtube & Instagram. Ranking based on likes so the best jokes of the day are trending on top of the feed. Recommendation engine with a multibandid ML algo from the start so you can leverage all that incoming data.
- sunsetMurk 6y agoAwesome. here come the 'deepMemes'!
- echelon 6y ago> Can you give some more info on how you generated the models? glow-tts and melgan, which are somewhat unpopular choices given the proliferation of Tacotron2/Waveglow. I chose these due to their sparsity and speed. > I'm also interested in the tech stack you're using to implement this webapp... Would love some details! It's a Rust microservice architecture. There's a proxy layer that decodes the request and sends it to the appropriate backend, and then there's the tts service that is horizontally scaled and is responsible for loading the model pipeline and turning requests into audio. > ..What's next? For me? Voice conversion in the near term. This takes microphone input and turns it into the target speaker's voice. I'm also spending a lot of time on photogrammetry. I have a 3d volumetric webcam system right now that I have much bigger plans for.