6 ms·
I've been dying to replace my Echos with an open source smart speaker but half of them use AWS or Azure for test to speech and speech synthesis so really all yo
by throwaway2016a 4y ago
I've been dying to replace my Echos with an open source smart speaker but half of them use AWS or Azure for test to speech and speech synthesis so really all you are in control of is the software that runs on the device itself. So this is a coo step in the right direction.
- kelnos 4y agoThey're also just really not great. I tested out Mycroft a couple years ago and found that the success rate for getting it to understand its wake word and listen for commands was under 10%. Maybe if you buy their prepackaged product, it works better, but that's not something I want to do. I just want to run it on a Pi 4 (which they claim works) with a mic array.
- stavros 4y ago10% doesn't sound much worse than my Alexas' 30%...
- kevinmgranger 4y agoAnd if it's anything like siri, it can barely do anything useful, so it doesn't matter if it understands you.
- jjeaff 4y ago"I searched the web for 'shutup stop mute stop talking' and here is what I found on Wikipedia..."
- slickdork 4y agocounter point, I use my alexa daily and don't run into many issues with voice recognition, or it's lack of understanding. Daily uses: 1) in the AM i set up all my needed reminders for 5 minutes before every meeting I have 2) it's connected to my hue bridge so I can turn off/on lights by asking while laying in bed, which is wonderful. 3) I play music all day. 4) It reminds me 10 minutes before every sunset to go outside for a walk.
- 314 4y agoI find Siri really useful - for a very limited set of tasks where recognition is about 100% and being hands free has a benefit. Typically this is starting exercise workouts and countdown timers. For more general tasks the recognition is still good (for me, seems to cary by voice) but even at 90% there will be one mistakes in most requests.
- kelnos 4y agoMy Google home is pretty near 100%; I can count on one hand the number of times it hasn't "heard" me over the past year. That's my benchmark.
- Brendinooo 4y agoThey’ve done a lot of work in the last year on the software side. Might be worth revisiting. They’re tentatively on track to (finally!) ship in September of this year.
- krisgesling 4y agoYeah I think there are two sides to this coin (and just for clarity - all of this relates to Picroft, not Mimic 3 the TTS engine that just launched). The audio hardware makes a huge difference to audio input which is why we've developed the custom SJ201 board that's in the Mark II. But even on DIY units we have been making big improvements on the wake word detection by better balancing our training data sets. Once the Mark II is shipping there are additional wake word improvements on the roadmap. Eventually the system will optimize for the users of each device. So the wake word model on your device wouldn't be exactly the same as the model on mine. We've also ported the Wake Word model to Tensorflow Lite which means it uses a small fraction of the system resources that it used to :D We're also about to make some bigger changes to mycroft-core that will help to support a broader range of hardware in a more consistent way. So whilst you could try it again today and I can guarantee it's better than the last time you used it, if you want a DIY system instead of a Mark II - I'd suggest adding a reminder to check it again in a couple of months once these bigger changes land.
- kelnos 4y agoThat sounds fantastic! Thank you for replying; I'll definitely check back and give it another go.
- Semaphor 4y agoThe Rhasspy [0] author recently got hired by mycroft to work on satelites and fully local. Rhasspy requires a lot of manual work, but replacing Alexa is already possible. I’m somewhat stuck with the current hardware availability issues, but I have a Pi 3 satellite that does wakeword detection (this is supposed to be handled by Pi Zero 2 W in the future) and sends the voice to the MQTT server running on a PI 4, the data gets picked up by the Rhasspy instance also running there, it does STT, intent recognition, sends the intent to home assistant and then does TTS back to the satellite. My main software issue is currently how to replicate the music functionality. Playing music at the satellite that requested it, lowering the volume when it recognizes the wakeword. Preselection of "commands" for band and genre names should be easily scriptable afterwards. In a quiet room, I have no issues with wakeword detection using a playstation eye camera (I wanted the seed USB microhphone array, but between discovering it and starting with buying hardware the supply chain bit once again) [0]: https://rhasspy.readthedocs.io/en/latest/ https://rhasspy.readthedocs.io/en/latest/
- puchatek 4y agoAnd how well does STT work when the room is not quiet anymore, e.g. when music is playing?
- Semaphor 4y agoMy understanding is, that the seeed array would work better than the PS eye, but for the volume I normally listen music at, it still works okay.
- krisgesling 4y agoYeah we aren't using the seeed array in the final Mark II. But we have used the same XMOS XVF-3510 to perform acoustic echo cancellation. That means, even with music blasting out of the speakers, you can still wake the device from across the room. In a simple fashion you can think of it as subtracting the audio being output from the audio coming in from the microphone.