8 ms·
AudioGen: Textually Guided Audio Generation
- fuzzythinker 4y ago[code] redirects to the same page
- ggerganov 4y agoAccording to one of the authors, the code and the models will be available soon [0] [0] - https://twitter.com/FelixKreuk/status/1575846953333579776 https://twitter.com/FelixKreuk/status/1575846953333579776
- nudpiedo 4y agoThat could be another missing piece to videogame generational art, sfx sounds and soon soundtracks.
- deleted 4y ago[deleted]
- karmasimida 4y agoIt will be more useful if it can narrate text along with those background effects.
- simonw 4y agoYou can already achieve that by combining models - use a dedicated speech synthesis model for the narration, then layer that over background effects from AudioGen. Given that, I don't think AudioGen particularly needs to add full narration. That seems like a very different problem to me, likely requiring a completely different architecture.
- godmode2019 4y agoWhat is the current state of the art speech synthesis model?
- ricopags 4y agoIt was Nvidia's Tacotron2[0] but now I believe it's NaturalSpeech[1] [0]https://paperswithcode.com/method/tacotron-2 https://paperswithcode.com/method/tacotron-2 [1]https://speechresearch.github.io/naturalspeech/ https://speechresearch.github.io/naturalspeech/
- godmode2019 4y agoThank you
- kevmo314 4y agoThe speech samples are really funny. Very Sims-esque.
- solardev 4y agoThe last thing you'll hear before the AI eats you: https://felixkreuk.github.io/text2audio_arxiv_samples/large_32factor_1streams_2048codesPerBook/continuous_laughter_and_chuckling.mp3 https://felixkreuk.github.io/text2audio_arxiv_samples/large_...
- uwagar 4y agos/textually/sexually i giggled :)
- iamthemonster 4y agoIt would be very interesting indeed to have an ebook reader paired with bluetooth earphones, and it simultaneously feeds the words into this to make an ambient soundtrack, perhaps also choosing music appropriate to the word-choice on the page.
- youssefabdelm 4y ago-__- I wish researchers would train a stereo 44.1kHz version...why always 16kHz? I know I know 16kHz saves more compute but come ooooon you're Meta
- fragmede 4y agoText2audio is impressive, but I wanna see dance2audio. Just need a million dollars in funding to pay for cameras and dancers.
- creative2022 4y ago
- creative2022 4y ago