4 ms·
Sadly looks unlikely if the base model wasn't trained on vocals. > Mitigations: Vocals have been removed from the data source using corresponding tags, and the
by operator-name 3y ago
Sadly looks unlikely if the base model wasn't trained on vocals.
> Mitigations: Vocals have been removed from the data source using corresponding tags, and then using a state-of-the-art music source separation method, namely using the open source Hybrid Transformer for Music Source Separation (HT-Demucs).
> Limitations: The model is not able to generate realistic vocals.
(https://github.com/facebookresearch/audiocraft/blob/main/model_cards/MUSICGEN_MODEL_CARD.md https://github.com/facebookresearch/audiocraft/blob/main/mod...)
I suspect this was a combination of playing it safe and that the model isn't well architected to reproduce meaningful vocals.