19 ms·
The era of open voice assistants
- aaron695 2y ago[dead]
- AIFounder 2y ago[dead]
- jfim 2y agoThat's a pretty timely release considering Alexa and the Google assistant devices seem to have plateaued or are on the decline.
- IgorPartola 2y agoCurious what you mean by that.
- oaththrowaway 2y agoFor me the Alexa devices I own have gotten worse. Can't do simple things (setting a timer used to be instant, now it takes 10-15 seconds of thinking assuming it heard properly), playing music is a joke (will try to play through Deezer even though I disaled that integration months ago, and then will default to Amazon Music instead of Spotify which is set as the default). And then even simple skills can't understand what I'm asking 60% of the time. The first maybe 2 years after launch it seemed like everything worked pretty good but since then it's been a frustrating decline. Currently they are relagated to timers and music, and it can't even manage those half the time anymore.
- interludead 2y agoThat aligns with some of the frustration I’ve heard from others. It’s surprising (and disappointing) how these platforms, which seemed to have so much potential early on, have started to feel more like a liability
- lelag 2y agoIt is, I think, a common feeling among Echo/Alexa users. Now that people are getting used to the amazing understanding capabilities of ChatGPT and the likes, it probably increases the frustration level because you get a hint of how good it could be. I believe it boils down to two main issues: - The narrow AI systems used for intent inference have not scaled with the product features. - Amazon is stuck and can't significantly improve it using general AI due to costs. The first point is that the speech-to-intent algorithms currently in production are quite basic, likely based on the state of the art from 2013. Initially, there were few features available, so the device was fairly effective at inferring what you wanted from a limited set of possibilities. Over time, Amazon introduced more and more features to choose from, but the devices didn't get any smarter. As a result, mismatches between actual intent and inferred intent became more common, giving the impression that the device is getting dumber. In truth, it’s probably getting somewhat smarter, but not enough to compensate for the increasing complexity over time. The second point is that, clearly, it would be relatively straightforward to create a much smarter Alexa: simply delegate the intent detection to an LLM. However, Amazon can’t do that. By 2019, there were already over 100 million Alexa devices in circulation, and it’s reasonable to assume that number has at least doubled by now. These devices are likely sold at a low margin, and the service is free. If you start requiring GPUs to process millions of daily requests, you would need an enormous, costly infrastructure, which is probably impossible to justify financially—and perhaps even infeasible given the sheer scale of the product. My prediction is that Amazon cannot save the product, and it will die a slow death. It will probably keep working for years but will likely be relegated by most users to a "dumb" device capable of little more than setting alarms, timers, and providing weather reports. If you want Jarvis-like intelligence to control your home automation system, the vision of a local assistant using local AI on an efficient GPU, as presented by HA, is the one with the most chance of succeeding. Beyond the privacy benefits of processing everything locally, the primary reason this approach may become common is that it scales linearly with the installation. If you had a cloud-based solution using Echo-like devices, the problem is that you’d need to scale your cloud infrastructure as you sell more devices. If the service is good, this could become a major challenge. In contrast, if you sell an expensive box with an integrated GPU that does everything locally, you deploy the infrastructure as you sell the product. This eliminates scaling issues and the risks of growing too fast.
- IgorPartola 2y agoThat’s interesting because I have a bunch of Echos of various types in my house and my timers and answers are instant. Is it possible your internet connection is wonky or you have a slow DNS server or congested Wi-Fi? I don’t have the absolute newest devices but the one in my bedroom is the very original Echo that I got during their preview stage, the one in my kitchen is the Echo Show 7” and I have a bunch of puck ones and spherical ones (don’t remember the generations) around the house. One did die at one point after years of use and got replaced but it was in my kids room so I suspect it was subject to some abuse.
- creeble 2y agoI too get pretty consistent response and answers from Alexa these days. There has been some vague decline in the quality of answers (I think sometime back they removed the ability to ask for Wikipedia data), but have no trouble with timers and the few linked wemo switches I have. I’m also the author of an Alexa skill for a music player (basic “transport” control mostly) that i use every day, and it still works the same as it always did. Occasionally I’ll get some freakout answer or abject failure to reply, but it’s fairly rare. I did notice it was down for a whole weekend once; that’s surely related to staffing or priorities.
- mrweasel 2y agoAmazon also fired a large number of people from the Alexa team last year. I don't really think Alexa is a major priority for Amazon at this point. I don't blame them, sure there are millions of devices out there, but some people might own five device. So there aren't as many users as there are devices and they aren't making them any money once bought, not like the Kindle. Frankly I know shockingly few people who uses Siri/Alexa/Google Assistant/Bixby. It's not that voice assistants don't have a use, be it is a much much small use case than initially envisioned and there's no longer the money to found the development, the funds went into blockchain and LLMs. Partly the decline is because it's not as natural an interface as we expected, secondly: to be actually useful, the assistants need access to control things that we may not be comfortable with, or which may pose a liability to the manufacturers.
- bdavbdav 2y agoGH is basically abandonware at this stage it seems. They just seem to break random things, and there haven’t been any major updates / features for ages (and Gemini is still a way off for most).
- cachvico 2y agoGoogle Home's Nest integration is recent and top-notch though. Hopefully in a year they'll have rolled out the Gemini integration and things will be back on track.
- bdavbdav 2y agoI’d not go as far as too notch. We’ve reverted as family members don’t get notifications (like from the doorbell)
- lolinder 2y agoOn the Google side it's become basically useless for anything beyond interacting with local devices and setting timers and reminders (in other words, the things that FOSS should be able to do very easily). Its only edge over other options used to be answering questions quickly without having to pull out a screen, but now it refuses to answer anything (likely because Google Search has removed their old quick answers in favor of Gemini answers).
- stickfigure 2y agoI was an early adopter of google home, have had several generations (including the latest). I quite like the devices, but the voice recognition seems to be getting worse not better. And the Pandora integration crashes frequently. In addition, it's a moron. I'm not sure it's actually gotten dumber, but in the age of chatgpt, asking google assistant for information is worse than asking my 2nd grader. Maybe it will be able to quote part of a relevant web page, but half the time it screws that up. I just want it to convert my voice to text, submit it to chatgpt or claude, and read the response back to me. All that said, the audio quality is good and it shows pictures of my kid when idle. If they suddenly disappeared I would replace them.
- throwawayq3423 2y agoGoogle and Amazon refuse to put GenAI into their existing speakers (which barely function). No doubt they want a new product launch to charge more.
- frognumber 2y agoI don't fully understand the cloud upsell. I have a beefy GPU. I would like to run the "more advanced" models locally. By "I don't fully understand," I mean just that. There's a lot of marketing copy, but there's a lot I'd like to understand better before plopping down $$$ for a unit. The answers might be reasonable. Ideally, I'd be able to experiment with a headset first, and if it works well, upgrade to the $59 unit. I'd love to just have a README, with a getting started tutorial, play, and then upgrade if it does what I want. Again: None of this is a complaint. I assume much of this is coming once we're past preview addition, or is perhaps there and my search skills are failing me.
- trb 2y agoFinding microphones that look nice, can pick up voice at high enough quality to extract commands and that cover an entire room is surprisingly hard. If this device delivers on audio quality it's totally worth it at $59.
- bdavbdav 2y ago100%. For a lot of users that have WAF and time available to contend with, this is a steal. Bear in mind that a $50 google home or Alexa mini(?) is always going to be whatever google deem it to be. This is an open device which can be whatever you want it to be. That’s a lot of value in my eyes.
- alias_neo 2y agoI've found it quite hard to find decent hardware with both the input capability needed for wakeword and audio capture at a distance, whilst also having decent speaker quality for music playback. I started using the Box-3 with heywillow which did amazing input and processing using ML on my GPU, but the speaker is aweful. I build a speaker of my own using a raspberry pi Z2W, dac and some speakers in a 3d printed enclosure I designed, and added a shim to the server so that responses came from my speaker rather than the cheap/tiny speaker in the box-3. I'll likely do the same now with the Voice PE, but I'm hoping that the grove connector can be used to plonk it on top of a higher quality speaker unit and make it into a proper music player too. As soon as I have it in my hands, I intend to get straight to work looking at a way to modify my speaker design to become an addon "module" for the PE.
- thumbsup-_- 2y agoWe need more projects like home assistant. I started using it recently and was amazed. They sell their own hardware but the whole setup is designed to works on any other hardware. There are detailed docs for installation on your own hardware. And, it works amazingly well. Same for their voice assistant. You can but their hardware and get started right away or you can place your own mics and speakers around home and it will still work. You can but your own beefy hardware and run your own LLM. The possibilities with home assistant are endless. Thanks to this community for breaking the barriers created by big tech
- mkagenius 2y agoI am working on automation of phones (open source) - https://github.com/BandarLabs/clickclickclick https://github.com/BandarLabs/clickclickclick I haven't been able to quite get the Llama vision models working but I suppose with new releases in future, it should work as good as Gemini in finding bounding boxes of UI elements.
- lokar 2y agoIt’s a great project overall, but I’ve been frustrated by how anti-engineer it has been trending.
- thfuran 2y agoHow so?
- snailmailman 2y agoIm a different user- but I can say I’ve been frustrated with their refusal to support OIDC/oauth/literally any standard login system. There is a very long thread on their forums documenting the many attempts for people to contribute this feature.[0] The devs simply shut it down every time, with little to no explanation. I run many self hosted applications on my local network. Homeassistant is the only one I’m running that has its own dedicated login. Everything else I’m using has OIDC support, or I can at least unobtrusively stick a reverse proxy in front to require OIDC login. [0] https://community.home-assistant.io/t/open-letter-for-improving-home-assistants-authentication-system-oidc-sso/494223 https://community.home-assistant.io/t/open-letter-for-improv... Edit: things like this [1] don’t help either. Where one of the HA devs threatens to relicense a dependency so that NixOS can’t use it, because… he doesn’t want them to? The license permits them to. Seemed very against the spirit of open source to me. [1] https://news.ycombinator.com/item?id=27505277 https://news.ycombinator.com/item?id=27505277
- Jarwain 2y agoI'm actually really excited for this! I noticed recently there weren't any good open source hardware projects for voice assistants with a focus on privacy. There's another project I've been thinking about where I think the privacy aspect is Important, and figuring out a good hardware stack has been a Process. The project I want to work on isn't exactly a voice assistant, but same ultimate hardware requirements Something I'm kinda curious about: it sounds like they're planning on a sorta batch manufacturing by resellers type of model. Which I guess is pretty standard for hardware sales. But why not do a sorta "group buy" approach? I guess there's nothing stopping it from happening in conjunction I've had an idea floating around for a site that enables group buys for open source hardware (or 3d printed items), that also acts like or integrates with github wrt forking/remixing
- Brendinooo 2y agoI invested in Mycroft and it flopped. Here’s hoping some others can go where they couldn’t.
- deleted 2y ago[deleted]
- bdavbdav 2y agoI guess the difference here is that HA has a huge community already. I believe the estimate was around 250k installations running actively. I suspect a huge chunk of the HA users venn diagram slice fits within the voice users slice.
- balloob 2y agoOur estimates are more than a million active instances https://analytics.home-assistant.io/ https://analytics.home-assistant.io/
- emsixteen 2y agoMore than a million? It says on the page: "424,548 Active Home Assistant Installations" Am I missing something? Is it that these are just those you know are sharing details, and you can scale that up by a known percentage? :)
- jauntywundrkind 2y agoNot super convinced the XMOS audio processing chip is really gonna buy a lot. Trying to do audio input processing feels like a dynamic task, requiring such adaption. XMOS is the most well known audio processor and a beast, but not sure it's really gonna help here! I really hope we see some open-source machine -learned systems emerge. I saw Insta360 announce their video conferencing solution today. Optics looks pretty medium, nothing wild, but Insta360 is so good at video that I expect it'll be great. But there's a huge 14 microphone array on it, and that's the hard job; figuring out how to get good audio from speakers in a variety of locations around a room. It really made me wish for more open source footing here, some promising start, be it the conference room or open living space. I've given all of 60s to look through this, and was kinda hopeful because heck yeah Home Assistant, but my initial read isn't super promising, isn't that this is starting the proper software base needed to listen well to the world. https://petapixel.com/2024/12/17/the-insta360-connect-is-a-2000-videoconferencing-camera-on-steroids/ https://petapixel.com/2024/12/17/the-insta360-connect-is-a-2...
- choffee 2y agoThey showed a video at the end of their broadcast last night comparing what the raw microphone hears and what comes out of the XMOS chip and you can hear a much clearer voice all the time even when there is noise or you are far away from the device. It is also used to cancel out the music if you are using it's speaker output. I don't think it's doing any voice processing but it's cleaning up the audio a lot which makes the job of the wake word processor and the speach to text a lot easier. Up until now this was missing from a lot of the home made voice assistance and I think why Alexa can understand you from the next room but my home made one struggles with all but quiet conditions.
- summm 2y agoAlexa Echo Dot has 6 or 7 microphones. I'd expect that makes it much easier to filter out voices directionally than only the 2 microphone this hardware has. I hope they release a version with more microphones.
- nickthegreek 2y agoAnd on back order everywhere. I just spent the last 2 weeks getting a esp32-s3-box setup to do this but its lack of audio out really irks me.
- joshstrange 2y agoAnd the mic is not all that great either. I have a couple of them but they just weren't reliably picking up my voice and I couldn't hear the reply either (when it did hear me). I figured it would be easy to add a speaker to them but that sent me down a rabbit hole that I gave up on and put them in a drawer. I'll buy this for sure though because when the ESP32 box thing worked it worked really well and I loved being able to swap out parts of the assist pipeline.
- nickthegreek 2y agoI ended up moddng the s3 yaml to turn off the internal speaker and to forward all voice responses to a google hub.
- alias_neo 2y agoTo be fair, the issue with the Box-3 is HA's implementation; I used it with heywillow.io and it was incredible, I could speak to it from another room and it would pick up perfectly. The audio out is terrible so I wrote a shim-server that captures the request to the TTS server for heywillow and sent it to a speaker I build myself running MPD on a Pi with a nice DAC and have it play the responses instead of the box-3's tiny speaker. I don't expect the audio-out on this to be much better with its tiny speaker, but at least it has a 3.5mm jack. I'm going to look into what that Grove port can do too and perhaps build a new speaker "module" that the Voice PE can sit on top of to make it a proper music device.
- yzydserd 2y ago> And on back order everywhere. I just clicked through to my large country and the first vendor and was able to buy 2 for delivery tomorrow. So it says. So maybe not on back order everywhere.
- 2y ago
- mkagenius 2y agoThough a separate hardware helps - I believe voice and automation can be integrated more seamlessly to our existing devices (phones/laptops) with high compute built in. Llama and whisper are already public so that should help innovation in this area.
- antonyt 2y agoWith existing phones and laptops, there’s either activation friction (pressing the “listen to me” button) or the device has to be always listening, which requires a lot of trust in your hardware vendors. With an open source and potentially local-only device, you can have your voice assistant and keep your privacy.
- throwawaymaths 2y agolast i checked open source whisper does not support streaming or diarization out of the box. you really need both for a good voice assistant experience
- joshstrange 2y agoYou can use your phone to text or talk to HA's assistant. I've done that a number of times when Alexa fails. Having dedicated hardware is a huge step up for me. I've tried their ESP32 mini cube assistant thing before and it showed a lot of promise but the hardware (speaker and mic, processor was fine) was lacking. This seems to be a good mic and speaker wrapped around a similar core so I'm super excited for it.
- alias_neo 2y agoThe voice input can really be done however you like, the benefit of a device like the Voice PE is the wake word detection on-device. I have an office-style desk-phone (SNOM) connected to a SIP server and I can pick the receiver up and talk to the Assistant, but you can plug in any way you like to get the audio to/from HA. With your phone, wake words are usually locked down by Apple/Google so you can't really have it hands-free, and that's the problem this device is solving; not the audio input itself, but the wake-word/handfree input. On an Android phone, you can replace the Google Assistant with the Home Assistant one, but you still have to activate it the usual way, press a button or launch the app etc.
- shepherdjerred 2y agoHome Assistant is such a fantastic project. I've been waiting for something like this for a long time; I just pre-ordered three. My only remaining wish is that I can replace Siri with this (without needing some workaround)
- deleted 2y ago[deleted]
- hoppp 2y agoIf it runs fully on premise that would be great. Im still not comfortable buying a device that records everything I say and uploads it to a cloud
- haddonist 2y agoFully on-prem can be done if you've got the LLM compute power in place.
- catmanjan 2y agoAll I want is a voice assistant that I can call "computer" like Star Trek, I don't want to have to say a brand name thankyou!
- dartos 2y agoYou could’ve always set Alexa to respond to “Computer” instead.
- catmanjan 2y agoAh I admit I haven't looked into it for several years, good to see they added the feature - I might have to grab one
- bigstrat2003 2y agoThe problem is that it will go off every single time you watch Star Trek.
- imp0cat 2y agoCan confirm, this works fabulously!
- antonyt 2y agoIf you run openWakeWord, “computer” is one of very many pretrained models the community has made: https://github.com/fwartner/home-assistant-wakewords-collection/tree/main/en https://github.com/fwartner/home-assistant-wakewords-collect...
- lxe 2y agoHere's what I'm looking for in a voice assistant: - Full privacy: nothing goes to the "cloud" - Non-shitty microphones and processing: i want to be able to be heard without having to yell, repeat, or correct - No wake words: it should listen to everything, process it, and understand when it's being addressed. Since everything is private and local, this is now doable - Conversational: it should understand when I finished talking, have ability to be interrupted, all with low latency - Non-stupid: it's 2024, and alexa and siri and google are somehow absolutely abysmal at doing even the basics - Complete: i don't want to use an app to get stuff configured. I want everything to be controlled via voice
- nissarup 2y agoLooks like you are in the market for a butler. Especially your last point will, IMO, not be possible for a long time.
- danparsonson 2y ago> No wake words: it should listen to everything, process it, and understand when it's being addressed Even humans struggle with this one - that's what names are for!
- antonyt 2y agoYeah, I’m having a hard time imagining how no-wake-word could work in practice.
- fragmede 2y agoafter setting up the system, if I say "turn the ceiling lights to 20%", who else would be changing the lights? But also, post-fix wake word would also be natural if it was recording all the time. "turn on the lights, Google", for instance
- TheCoelacanth 2y agoSomeone in a TV show that you're watching?
- IG_Semmelweiss 2y agoCan someone describe the use case here? I don't quite understand what its purpose is. Is this a fully-private, open source alternative to Alexa, that by definition requires a CPU locally to run ? Is the device supposed to be the nerve center of IoT devices ? Can it access the Wifi to do web crawls on command (music, google, etc)?
- IvyMike 2y agoIf you have home automation, surely you've run into this situation when Comcast flakes (or similar): "OK, Google, turn lights on" "Check your connection and try again" As far as I can tell, if you have Home Assistant + this new device, you've fixed that problem.
- antonyt 2y agoThe nerve center would be your Home Assistant instance, which is not this device. You can run Home Assistant on whatever hardware you like, including options sold by Nabu Casa. This device provides the microphone, speaker, and WiFi to do wake-word detection, capture your input, send it off to your HA instance, and reply to you with HA’s processed response. Whether your HA instance phones out to the internet to produce the response is up to you and how you’ve configured it.
- deleted 2y ago[deleted]
- jve 2y agoWhile we are getting shoveled AI keyword everywhere, I'm actually disappointed I don't see it here. The first thought I had when encountering LLM was that it can finally make these devices understand you and make them finally useful... and I don't need to know some presceipted keywords.
- antonyt 2y agoYou can actually integrate LLMs with Assist pipelines, it’s just orthogonal to this hardware announcement. Check out https://www.home-assistant.io/blog/2024/06/05/release-20246/#dipping-our-toes-in-the-world-of-ai-using-llms https://www.home-assistant.io/blog/2024/06/05/release-20246/...
- pimeys 2y agoIt's also really cool. You can make it so that the home assistant itself first tries to understand what you do, like turning on the living room lights or setting the bathroom temperature to 21.5 degrees celsius. If the assistant pipeline does not understand what you are asking for, it can send your question to the LLM of your choice. You can also make the LLM to control the lights, heat etc, but at least for now ChatGPT is pretty bad with that. So let home assistant do the home automation, and then let ChatGPT to answer your questions about the most popular ruler in the 19th century France.
- joshstrange 2y agoIt's too bad it's sold out everywhere. I've tried the ESP32 projects (little cube guy) for voice assistants in HA but it's mic/speaker weren't good enough. When it did hear me (and I heard it) it did an amazing job. For the first time I talked to a voice assistant that understood "Turn off office lights" to mean "Turn off all the lights in the office" without me giving it any special grouping (like I have to do in Alexa and then it randomly breaks). It handled a ton of requests that are easy for any human but Alexa/Siri trip up on. I cannot wait to buy 5 or more of these to replace Alexa. HA is the brain of my house and up till now Alexa provided the best hardware to interact with HA (IMHO) but I'd love something first-party.
- bdavbdav 2y agoHow did you find it for music tasks?
- joshstrange 2y agoI didn’t test that. I normally just manually play through my Sonos speaker groups on my phone. I don’t like the sound from the Echos so I’m not in the habit of asking them to do anything related to music. Right now I only use Alexa for smart house control and setting timers
- moffkalast 2y agoI'm definitely buying one for robotics, having a dedicated unit for both STT and TTS that actually works and integrates well would make a lot of social robots more usable and far easier to set up and maintain. Hopefully there's a ROS driver for it eventually too.
- shaklee3 2y agoAs someone not that familiar with haas, can someone explain why there's not a clear path to replace Alexa or Google home? I considered using haas recently to get a gpt like response after being frustrated with Google home, but it seems this is a complete mess. is there a way to get this yet?
- joshstrange 2y ago> explain why there's not a clear path to replace Alexa or Google home? There is. I've used HA with their default assist pipeline (Cloud HA STT, Cloud HA LLM, Cloud HA TTS) and I've also plugged in different providers at each step (both remote and local for each part: STT/LLM/TTS) and it's super cool. Their default LLM isn't great but it works, plugging in OpenAI made it work way better. My local models weren't great in speed but I don't have hardware dedicated for this purpose (currently), seeing an entire local pipeline was amazing for the promise of it in the future. It's too slow (on my hardware) but we are so close to local models (SST/TTS could be improved as well but they are much easier to do already locally). If this new HA hardware comes even close to performing as well as the Echo's in my house (low bar) I'll replace them all.
- jazzyjackson 2y agoWhat does it use LLMs for?
- joshstrange 2y agoTaking the text of what you said and figuring out what you want to do. It sends what you said plus a list of devices/states and a list of functions (to turn off/on, set temp, etc of devices). The LLM takes "Turn off basement lights" and turns that into "{function: "call_service", args: ['lights.on', 'entity-id-123']}" (<- Completely made up but it's something like that) that it passes back to HA along with what to say back to the user ("Lights turned off" or whatever) and HA will run the function and then do TTS to respond to you.
- tomqueue 2y agoI am very excited for this. One question I couldn’t find an answer for though is whether the hardware is open enough to be usable with other home automation systems. I am using OpenHAB and they too have an integrated voice assistant. I looked into migrating to HA a couple times but eventually gave up, primarily because it felt like such a waste of time to migrate a fully working environment with dozens of rules and scripts to yaml files.
- interludead 2y agoMoving a fully functional setup with complex rules and scripts is a daunting task
- choffee 2y agoIt's all open and so should be able to work with OpenHAB as well but it would need somebody to either write a firmware that's compatibale with the OpenHAB endpoints or add ESPHome interegeation into OpenHAB. Somebody might have already done that for their voice stuff. There is not much yaml in home assistant now unless you want it. I'd give it a go in a VM and see what it finds on your network :)
- interludead 2y agoI think in some ways it could redefine how we think about voice control... taking it from the cloud and putting it back into users' hands, like literally
- Simon_O_Rourke 2y agoA good emphasis in the summary, that certain other companies will only focus on monetization at the expense of features and functionality.
- delijati 2y agoPerfect will dig more into it. Currently i like to have an spotify client without ui for my kids ;)
- ahaucnx 2y agoIt's not clear to me from the description if this is also completely open source hardware. Are the schematics, BoM, firmware published under a permissible license? If so, where are they accessible? And if not, I would be curious to know why it haven't been fully open sourced.
- choffee 2y agoI would think so in the end. They talked about the case design being open. The software and firmware are all open already and they said that they really wanted people to be able to take these components and make new devices. They have relesased the designs for the yellow so I assume it will all come. https://github.com/NabuCasa/yellow https://github.com/NabuCasa/yellow
- leeoniya 2y agoanyone tried https://getleon.ai/ https://getleon.ai/ ?
- lukifer 2y agoI tried years ago, I don't think I got it working, ended up using Rhasspy/voice2json instead (TIL: the creator of both is now the Voice Eng Lead for Home Assistant). Looks like the GitHub is still somewhat active, although their roadmap links to a dead Trello: https://github.com/leon-ai/leon https://github.com/leon-ai/leon
- fx1994 2y agoWhat I don't like is that most voice assistances perform really bad on my native language so I don't use them at all. For english speakers yes, but for all other not so much. I guess it will get better.
- choffee 2y agoThat is one of the major things that Home Assistant are trying to fix. They have groups working on most languages and are adding them to their open as they improve. https://www.home-assistant.io/voice_control/contribute-voice https://www.home-assistant.io/voice_control/contribute-voice
- solarkraft 2y agoRIP Mycroft. A tad too early.
- choffee 2y agoNabu Casa employ one of the Mycroft devs now and i think some of the tech came from that project so it's not all gone :)
- singularity2001 2y agosorry if this question takes away from the great strives the team went through but wouldn't it be much easier (hardware wise) to jailbreak one of the existing great hardware thingies like Apple HomePod or the Google one or Alexa?
- choffee 2y agoI don't think they are that easy to jail break but I may be wrong. I think they wanted to create an open device that people could build from rather than just a hacked up alexa.
- robotfelix 2y agoI've picked up an Echo Dot a few years ago when Amazon were practically giving them away, thinking that surely someone would have jailbroken it by now to allow it to be used with Home Assistant. It was only after researching later that I discovered that this wasn't currently possible and recommended approach was to buy some replacement internals that cost more than the device itself (and if I recall correctly, more than the new Home Assistant Voice Preview Edition).
- alias_neo 2y agoThe fact that it hasn't (widely?) been done yet suggests the answer is "no". The hardware in those devices is generally better, most of them have much better speakers, but they're locked down, the wake-word detection hardware isn't open or accessible so changing it to do what we need would be difficult, and you're just hoping there's a way in. Existing examples of opening them (as in freedom) replace the PCB entirely, which puts you back to square one of needing open hardware. This feels like the right approach to me; I've been building my own devices for this purpose with off-the-shelf parts, and designing enclosures, but this is much sleeker; I just hope an add-on or future version comes with much better audio out (speakers) because that's where it and things like it (e.g. the S3-Box-3) are really lacking.
- singularity2001 2y agoor maybe find cheap Chinese smart speaker which is hackable?
- Havoc 2y agoHad to laugh a bit at the caveat about powerful hardware. Was bracing myself for GPU and then it says N100 lol
- moooo99 2y agoI mean, comparatively many people are hosting their home Assistant on an raspberry Pi so it is relatively powerful :D
- geerlingguy 2y agoAnd the CM5 is nearly equivalent in terms of the small models you run. Latency is nearly the same, though you can get a little more fancy if you have an N100 system with more RAM, and "unlocked" thermals (many N100 systems cap the power draw because they don't have the thermal capacity to run the chip at max turbo).
- moffkalast 2y agoIf we're being fair you can more like, walk models, not run them :) An 125H box may be three times the price of an N100 box, but the power draw is about the same (6W idle, 28W max, with turbo off anyway) and with the Arc iGPU the prompt processing is in the hundreds, so near instant replies to longer queries are doable.
- fons 2y agoI wonder how this compares to the Respeaker 2 https://wiki.seeedstudio.com/ReSpeaker_Mic_Array_v2.0/ https://wiki.seeedstudio.com/ReSpeaker_Mic_Array_v2.0/ The respeaker has 4 mics and can easily cancel out the noise introduced by a custom external speaker
- stavros 2y agoI don't just want the hardware, I want the software too. I want something that will do STT on my speech, send the text to an API endpoint I control, and be able to either speak the text I give it, or live stream an audio response to the speakers. That's the part I can't do on my own, and then I'll take care of the LLMs myself.
- alias_neo 2y agoAll of these components are available separately or as add-ons for Home Assistant. I currently do STT with heywillow[0] and an S3-Box-3 which uses an LLM running on a server I have to do incredibly fast, incredibly accurate STT. It uses Coqui XTTS for TTS, with very high quality LLM based voice; you can also clone a voice by supplying it with a few seconds of audio (I tested cloning my own with frightening results). Playback to a decent speaker can be done in a bunch of ways; I wrote a shim that captures the TTS request to Coqui and forwards it to a Pi based speaker I built, running MPD which then requests the audio from the STT server (Coqui) and plays it back on my higher quality speaker than the crappy ones built in to the voice-input devices. If you just want to use what's available HA, there's all of the Wyoming stuff, openWakeword (not necessary if you're using this new Voice PE because it does on-device wakeword), Piper for TTS, or MaryTTS (or others) and Whisper (faster-whisper) for STT, or hook in something else you want to use. You can additionally use the Ollama integration to hook it into an Ollama model running on higher end hardware for proper LLM based reasoning. [0]heywillow.io
- stavros 2y agoI do the same, Willow has been unmaintained for close to a year, and calling it "incredibly fast" and "incredibly accurate" tells me that we have very different experiences.
- cranberryturkey 2y agohttps://linuxvoice.ai https://linuxvoice.ai
- hamilyon2 2y agoI had great trouble simply connecting Bluetooth speaker to use it as voice input and for sound output. The overall state of sound subsystem for diy voice assistant feels third-class at best.
- albybisy 2y agoi don't wanna talk to a computer
- cheema33 2y ago> i don't wanna talk to a computer You are in luck. You can get a human butler. But not for $59.
- starlite-5008 2y ago[dead]
- bsdice 2y agoMajel Barrett voice please.
- ragmondo 2y agoMy plea / request : Make a home assistant a DROP IN replacement for a standard light switch. It has power, its adds functionality from the get-go (smart lighting), it’s placed in a convenient position for the room and no extra wires etc required.
- Carrok 2y agoLook at Shelly light switches.
- NegativeK 2y agoAgreed. They sell UL rated models, have an option for cloud connectivity but zero requirement, your switch still works if the Shelly loses connectivity with whatever home automation server you have, and it's a small box that you wire in behind the switch.
- Carrok 2y agoThey also make drop in replacement dimmer switches. Even easier than the small box style. https://us.shelly.com/products/shelly-plus-wall-dimmer https://us.shelly.com/products/shelly-plus-wall-dimmer
- timdiggerm 2y agoYou've misunderstood what they're asking for. They're asking for Home Assistant hardware (microphone, speaker, wifi) that, instead of being a standalone box taking up space on the counter/table/etc, fits into the hole in the hall where they currently have a lightswitch.
- Carrok 2y agoI guess I did misunderstand, because that request seems strange to me. I’m assuming they have more than one switch. Which one should have Home Assistant on it? Seems like an odd deployment strategy. A pi isn’t that big..
- bradly 2y agoAre there any MacOS software versions of this? I've been looking for opensource wake-work for a "Hey Siri"-like integration, but I'm very apprehensive of anything, malicious or not, monitoring the sound input for a specific word in an efficient way.
- silentOpen 2y agoOpenWakeWord has worked well for me especially using well-trained models like “Hey, Mycroft”.
- unshavedyak 2y agoWell shoot. Now i want to record everything in my house and transcribe it for logs. I already wanted to do that but didn't think there was a sane way.. assuming this lets me create a custom pipeline, that's wicked
- dboreham 2y agoIt isn't even one year since the press stories about how dumb a product Alexa was and how it makes no money and all the devs are getting laid off. Something changed now?
- eightysixfour 2y agoIt was a bad product at making money for Amazon, but they are useful for smart homes. Home Assistant is pretty squarely in the smart home category. I bought two the second they were announced, I already use the software stack with the m5 atoms and they are terrible devices, but the software works well enough for me.
- iamjackg 2y agoWell, the various Echo devices were allegedly built as loss leaders in the hope people would use them to make orders on Amazon. This is backed by the most active open source project on GitHub, which already has extensive support for voice pipelines both with and without LLMs, and is likely priced sensibly. A lot has changed in the open source ecosystem since commercial assistants were first launched. We have reliable open source wakeword detectors, and cheap/free LLMs can do the intent parsing, response generation, and even action calling.
- weird-eye-issue 2y agoHuh? Being able to do things like turn off lights or change the TV volume with your voice is actually quite a nice convenience
- marcosdumay 2y agoIf it's not clear, the Home Assistant business plan is different from the Amazon one for Alexa... and the Home Assistant open source project is even more different.
- sirtaj 2y agoI've been using the HA cloud voice assistant on my phone for the past few weeks, and it's such a great change from Alexa, because integrating new services and adding sentences is actually possible. Alexa, on the other hand, won't even allow a third party app to read its shopping list. It's no longer clear to me why Alexa even exists any more except as a kitchen timer.
- nailer 2y agoYou should talk to Sonos about partnering with them. They currently have a very limited Sonos voice assist, plus Google Voice and Alexa, but the latter two are limited pre-LLM assistants. I’m assuming they eventually want to create their own LLM and something privacy focused would be good match for their customers. I don’t know how they feel about open source though
- joshstrange 2y agoI would love so much if I could integrate Home Assistant into my Sonos devices. In fact, even with this box, I’d be interested in plugging into one of my Sonos for the output at least.
- skyde 2y agohow does this compare to ESP32-S3-BOX-3B ?
- sreejithr 2y agoGenuine question - How hackable is this? Can I have the voice commands redirected to my backend server where I can process it as I please?
- balloob 2y agoThis is Home Assistant. Everything is hackable. Inside Home Assistant the processing is delegated to integrations providing Speech-to-Text, command processing, Text-to-Speech. You can make custom integrations for all of them
- entropicdrifter 2y agoIt's fully open-source. I think the default use-case is to have the voice commands processed locally
- throwawayq3423 2y agoProbably as much as any other smart speaker without having to give your data away.
- zbrozek 2y agoIs anyone aware of an effort to repurpose Echo hardware to do HA voice?
- drdaeman 2y agoI've looked into this, and found nothing. One can surely repurpose the case and speakers, but the microphones are soldered on-board, and the board is not hackable and needs to go. To best of my awareness, there are no ways to load a custom firmware on a newer Echo device - they're locked down pretty tight.
- IshKebab 2y agoLooks great! The biggest issue I see is music. 90% of my use is "play some music" but none of the major streaming music providers offer APIs for obvious reasons. I'm not sure how you can get around that really.
- antonyt 2y agoTo do this in Home Assistant, you'd probably want to run Music Assistant and integrate it in. Looks like they manage to support some streaming providers, not entirely sure how: https://music-assistant.io/music-providers/ https://music-assistant.io/music-providers/ Getting it to play the right thing from voice commands is a bit of a rabbit hole: https://music-assistant.io/integration/voice/ https://music-assistant.io/integration/voice/
- steelframe 2y agoIf it's possible for the hardware to facilitate a use case, the employees working on the product will try to push the limits as far as they possibly can in order to manufacture interesting and challenging problems that will get them higher performance ratings and promotions. They will rationalize away privacy violations by appealing to their "good intentions" and their amazing ability to protect information from nefarious actors. In their minds they are working for "the good guys" who will surely "do the right thing." At various times in the past, the teams involved in such projects have at least prototyped extremely invasive features with those in-home devices. For example, one engineer I've visited with from a well-known in-home device manufacturer worked on classifiers that could distinguish between two people having sex and one person attacking another in audio captured passively by the microphones. As the corporate culture and leadership shifts over time I have marginal confidence that these prototypes will perpetually remain undeveloped or on-device only. Apple, for instance, has decided to send a significant amount of personal data to their "Private Cloud" and is taking the tactic of opening "enough" if its infrastructure for third-party audit to make an argument that the data they collect will only be used in a way that the user is aware and approves of. Maybe Apple can get something like that to a good enough state, at least for a time. However, they're inevitably normalizing the practice. I wonder how many competitors will be as equally disciplined in their implementations. So my takeaway is this: If there exists a pathway between a microphone and the Internet that you are not in 100% control over, it's not at all unreasonable to expect that anything and everything that microphone picks up at any time will be captured and stored by someone else. What happens with that audio will -- in general -- be kept out of your knowledge and control so long as there is insufficient regulatory oversight.
- comradesmith 2y agoOpen source
- gh02t 2y agoYeah, OP is comparing this to Google/Amazon/Apple/etc devices but this is being developed by the nonprofit that manages development on Home Assistant and in cooperation with their large community of users. It's a very different attitude driving development of voice remotes for Home Assistant vs. large corporations. They've been around for a while now and have a proven track record of being actual, serious advocates for data privacy and user autonomy. Maybe they won't be forever, but then this thing is open source. The whole point is that you control what these things do, and that you can run these things fully locally if you want with no internet access, and run your own custom software on them if that's what you want to do. This is a product for the Home Assistant community that will probably never turn much of a profit, nor do I expect it is intended to.
- gigel82 2y agoWhat is a good GPU to put in a home server that can run the TTS / STT and the local LLM required to make this shine? A 3090 is too expensive and power hungry. Maybe a 3060 12Gb? Is there anything in the "workstation" lineup that is more efficient especially since I don't need the video outs?
- afh1 2y agoMy experience with home assistance voice pipeline is nothing works and stt is terrible. I'll have to wait and see the reviews.
- ryukoposting 2y agoMy wife and I have been very happy with Home Assistant so far. The one thing we're missing is voice control, and until now it seemed like there just wasn't a clean solution for HA voice control. You were stuck doing some hobbyist shenanigans and hand-writing boatloads of YAML, or you were hooking up a HomeKit/Alexa which defeats the purpose of HA. This is a game-changer. They recommend an N100 in the blog post, but I might buy one anyway to see if my HA box's Celeron J3455 will do the job.
- Animats 2y agoNice. A totally local voice assistant. This makes sense for cars, where there's much local stuff to control. But for a home unit, what do you want to do that is entirely local? Turning the heat up and down gets boring after a while. If it does entertainment selection or shopping, it needs outside world connections. (Today's rant: I recently purchased a humidifier. It's just a little unit with a water tank, a water-softening filter, and an ultrasonic vaporizer. That part works fine. Then there are the controls. All this thing really needs is an on-off switch and a humidity knob, and maybe lights for power, humidification, and water tank empty. But no. It has five touch buttons and a round display about four inches across. The display is on even if the unit is off. Pressing the on/off button turns it on. If it's humidifying, there's a whole light show. The tank lights up purple. Swooping arcs of blue run up both edges of the round display. It's very impressive, especially in a dark bedroom. If you press and hold the second button for two seconds, about half the light show is suppressed. There are three fan speeds, and a button for that. Only the highest one will propel the water vapor high enough to avoid it hitting the floor and uselessly condensing before it mixes with the air. So that feature was not necessary. The display shows one number. It's usually the current humidity, but if you press the humidity set button, the number displayed becomes the setting, which is changed upwards by successive presses until it wraps around. After a few seconds, the display reverts to current humidity. Turning the unit off or removing the water tank resets all settings to the default. This is the low-end unit. The next step up comes with an IR remote. It's one way - the remote has buttons but no display. Since you have to be close to the display to use the buttons effectively, that doesn't help much. The step up after that is, inevitably, a cloud-based phone app. So this thing could potentially be interfaced to a voice assistant. That's only useful if there's enough information coming back from the device that the assistant software knows what the device is doing, and the assistant software understands that device status. If all it does is send remote button pushes, the result will be frustration. So you need some degree of intelligence at both ends - the end that talks to the human, and the end that talks to the device. If the user says "House, it's too dry in here", the assistant system needs to be able to check the status of the humidifier. Has power? Talking? On? Humidity setting reasonable? Fan running? Tank not empty? If it can't do that, it's part of the problem, not part of the solution.)
- meragrin_ 2y ago
- lostmsu 2y agoWhat voices do they use?
- lizzas 2y agoOpen as in 3d print files, rpi etc.? If so this is the project I am looking for!
- amluto 2y agoOne thing that makes me nervous: Home Assistant has an extremely weak security model. There is recent support for admin users, and that’s about it. I’m sort of okay with the users on an installation having effectively unrestricted access to all entities and actions. I’m much less okay with an LLM having this sort of access. An actually good product in this space IMO needs to be able to define specific sets of actions and allow agents to perform only the permitted actions.
- Ey7NFZ3P0nzAe 2y agoYou can already choose which entity to expose to the LLMs