3 ms·
I can see some fantastic uses for this in generating complex acoustic environments to layer over TTS or real recordings for speech-to-text model training. I wo
by blackkettle 3y ago
I can see some fantastic uses for this in generating complex acoustic environments to layer over TTS or real recordings for speech-to-text model training. I wonder if that is occupying some kind of gray-area. For example you have 1000hrs of clean speech from the librispeech corpus. It would be trivial to use this tool and available weights to generate background noise, environmental noise and the like, and then layer this with the clean speech to cheaply train a much more robust model. The environmental audio you create would never be directly shared or sold, but it would impact the overall quality of the STT model that you train from the combined results.