3 ms·
I wrote two articles on how to build custom voice assistants using just a Raspberry Pi and a microphone, one in 2019 https://blog.platypush.tech/article/Build-y
by blacklight 3y ago
I wrote two articles on how to build custom voice assistants using just a Raspberry Pi and a microphone, one in 2019 https://blog.platypush.tech/article/Build-your-customizable-voice-assistant-with-Platypush https://blog.platypush.tech/article/Build-your-customizable-... and one in 2020 https://blog.platypush.tech/article/Build-custom-voice-assistants https://blog.platypush.tech/article/Build-custom-voice-assis....
It's definitely doable and I still have my own custom assistants in the house. However, I had to get around with a Snowboy model for hotword detection (and Snowboy is now basically abandoned), Mozilla DeepSpeech model for speech-to-text (and that's quite heavy), and Mycroft's mimic3 text-to-speech model (and Mycroft is now basically bankrupt). Then writing the integration is relatively easy - I used Platypush, but it can definitely be done with Home Assistant and OpenHAB too.
Compared to 3-4 years ago, I think we're now in a state where the content is no longer the issue (just plug into a LLM, and all of your text requests will get an answer), nor integrations are a problem (just write a Platypush event hook on speech detected, and you can connect it to everything, no need for "Works with Google/Alexa" labels). Text-to-speech synthesis has also become cheap and ubiquitous.
But the hotword detection and speech-to-text models are still IMHO the bottleneck. Hotword detection is a field where you need a very small and lightweight model that only detects a specific word or phrase in a very reliable way. Snowboy was an amazing FOSS project - which also came with this cool idea of "crowd-funded models", where in order to download a model for a certain hotword you were first supposed to provide three audio tracks where you say that word in order to improve the model. But it's now discontinued because it cost the volunteers too much to run the infra.
And Mozilla DeepSpeech is a relatively good choice for general-purpose speech-to-text, but it's heavy (it takes 100% of the CPU when it runs on a Raspberry Pi) and it's mostly optimized for English - even support for other Western languages is patchy.
If there are other open-source alternatives that solve these problems, I'd be very happy to learn about them. Once these blockers are removed, there should be really no reason for anyone to feed their audio streams to Google or Amazon.
- hacker_gal 3y agoit's still expensive and not easy to build reliable models. thus open-source is not a sustainable option for maintainers. you have to have talented people and high quality data. talented people have opportunity cost, they can go to big tech with high 6 or 7figure TCs. some former mozilla people did. i'm not criticizing them, we gotta accept the cost of living is crazy. i care about my compensation too and i do not have energy left to contribute to foss. i salut those who can. you cannot compare donations vs. big tech TC unless you have a trust fund i guess. I don't have old money, so not sure about that part. snips.ai is acquired by sonos. google and sonos are still in patent wars. i'm not sure whether sonos will survive. it's not easy if big tech lawyers come after you, even you are kinda big. mycroft was a part of patent war too and the poor guy dedicated all of his energy instead of trying to build something. https://mycroft.ai/blog/huge-win-for-mycroft-at-the-patent-trial-and-appeal-board/ https://mycroft.ai/blog/huge-win-for-mycroft-at-the-patent-t... spokestack.io is no longer active. stability ai will probably go bankruptcy or aws or nvidia will acquire them: https://futurism.com/the-byte/stable-diffusion-stability-ai-risk-going-under https://futurism.com/the-byte/stable-diffusion-stability-ai-... picovoice.ai tries to do things differently, combining on-device speech processing and subscription based model. it's the only company i know that has good hotword detection and speech-to-text. their free option is enough for my project https://picovoice.ai/pricing/ https://picovoice.ai/pricing/ you may ask why nobody uses their tech to build an alternative voice assistant. again i dont think smart speakers are profitable enough. when you can buy an echo for $20 it's hard to find somebody to pay $20 per month. silicon shortage and logistics costs increased the boms. it's not easy to compete with amazon which can burn $10 billion. openai api probably feeds their audio streams to microsoft. otherwise why would nuance move to the cloud when they can process on-device, right after microsoft acquires them. to find options not feeding big tech, we gotta start paying for non-big tech.