3 ms·
This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent wa
by varunytoons 3y ago
This a fantastic new development in the AI Audio space!
However, it's quite disappointing that the model is closed sourced.
Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced!
https://github.com/QwenLM/Qwen-Audio https://github.com/QwenLM/Qwen-Audio
Does anyone have suggestions for how to integrate this into your tech stack via an internal API?
Interested to hear the varying thoughts on this. From what I softly understand is that the model weights have to be swapped or altered per se to be able to commercially reuse this. Correct me if I'm wrong.
- two_in_one 3y agoThanks for the link. License is clear: Researchers and developers are free to use the codes and model weights of both Qwen-Audio and Qwen-Audio-Chat. We also allow their commercial use. and important, if you have more than 100m active users:) 4. Restrictions If you are commercially using the Materials, and your product or service has more than 100 million monthly active users, You shall request a license from Us So, looks like it's absolutely fine to use, except for IT behemoth. As for how to use, API, I think. Interesting applications are possible. Like interactive mobile robots. Assistants for people with disabilities, both software and wearable. Interesting times... this will be called AI revolution probably. It's already not a joke, after several ups and downs.
- Zuiii 3y ago> Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio https://github.com/QwenLM/Qwen-Audio Openly distributed perhaps, but definitely not open source. The license appears as closed as meta (research-only with some leeway for other uses.) Do you have any truely open-source general audio generation models yet? I know about StyleTTS2, which is open source (MIT) and uncensored, but that model focuses on speech generation only. Having an non proprietary model like audiobox or Qwen-Audio would be really nice.
- TaylorAlexander 3y agoNormally I want basically everything to be open source. But as soon as I saw that audio restyling demo, I began to feel concerned they may be open sourcing this. The model can take a sample speakers voice, new text to speak, and also a description of a new location (like a cathedral with many echoes, or other background noises) and produce a convincing new audio sample. This technology will present serious challenges for the verification of covertly recorded audio. It will of course ultimately become widespread but I’m not inherently bothered by the idea of slowing down its release. Giving researchers extra time to examine possible detection techniques seems helpful to me.
- two_in_one 3y agothis technology is already exploited in the wild for scams. probably for years. it's too late to worry or try stop. the question is what can we do about it. and the first is to make people aware of it.