4 ms·
Interesting, this is yet another open source project that relies on a proprietary wake word model. Why are there no open source wake word engines like there ar
by skykooler 3y ago
Interesting, this is yet another open source project that relies on a proprietary wake word model. Why are there no open source wake word engines like there are for speech recognition?
- ktm5j 3y agoProbably because no one has built one yet.. feel free to do so.
- thesh4d0w 3y agoThe process of generating the wake word data is pretty onerous - https://docs.espressif.com/projects/esp-sr/en/latest/esp32s3/wake_word_engine/ESP_Wake_Words_Customization.html https://docs.espressif.com/projects/esp-sr/en/latest/esp32s3...
- kkielhofner 3y agoShort answer: it’s a tiny niche that is still very difficult, expensive, and time consuming. This isn’t Whisper or an LLM that has applications in practically anything. You’re not going to see OpenAI, Meta, etc release a microcontroller optimized wake implementation because literally no one cares (by comparison). Yet it requires a not-completely dissimilar level of effort. Someone else linked to the Espressif training dataset requirements. That said, wake word comes up more than I could have ever imagined. Everyone has an opinion and exactly zero of them have any idea what they are talking about (that I’ve encountered). It’s getting very, very old and I’ve all but given up engaging on it because it always ends with the classic “Well couldn’t you just…”. Yes, in the 15 seconds it took you to write that comment you came up with something no one in the field has ever thought of before. If you can create a completely open source wake implementation that gets even remotely close to the performance and reliability of those from Espressif, etc (while running on a microcontroller) we would be thrilled to use it. You will have created the first of its kind, and you’ll be famous! The more likely outcome (as you dismissively said) will be “yet another open source project” that goes in the graveyard of completely unusable open source wake word implementations. There are plenty. As you note - all of them. Let’s just say I won’t be waiting for it.
- jononor 3y agoIn case anyone with practical skills in ML+deep learning (but not in audio or embedded) wants to tackle this project, I am willing to mentor on those things. Microcontroller-level audio ML is my speciality. One can find me in the Sound of AI Slack. Ping me there and we can create a channel etc. It will be a multi-month endeavour though. Getting to PoC level is quite quick, but then getting robust performance in nearfield / high SNR cases across diverse background noise is a lot more work. And then tackling low SNR and far-field like with Alexa is yet another level. Model architecture wise the problem is rather well understood, there are several good papers available from ARM etc. Deployement infrastructure also pretty good these days, for example with Tensorflow Lite Micro. In addition to ML skills, the project would need someone that is good at organizing volunteer outreach, in order to build a good sized dataset. The Espressif docs are a reasonable spec for something quite good on the voice side. But then we would also need a good dataset of background noise. Of course those with embedded/microcontroller skills are also very welcome.