9 ms·
Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that the
by state_less 1y ago
Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more.
The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. What a turn of events!
- tedivm 1y agoThis is exactly what I do. I have two 3090s at home, with Qwen3 on it. This is tied into my Home Assistant install, and I use esp32 devices as voice satellites. It works shockingly well.
- tomatoman 1y agoSeems interesting setup, do you have it documented anywhere, thinking of building one!
- joshgachnang 1y agoOoo interesting, I'd love to hear more about the esp32's as voice satellites!
- sho_hn 1y agoI assume it's very similar to what Home Assistant's backing commercial entity Nabu Casa sells with the "Home Assistant Voice PE" device, which is also esp32-based. The code is open and uses the esphome framework so it's fairly easy to recreate on custom HW you have laying around.
- tedivm 1y agoFor the physical hardware I use the esp32-s3-box[1]. The esphome[2] suite has firmware you can flash to make the device work with HomeAssistant automatically. I have an esphome profile[3] I use, but I'm considering switching to this[4] profile instead. For the actual AI, I basically set up three docker containers: one for speech to text[5], one for text to speech[6], and then ollama[7] for the actual AI. After that it's just a matter of pointing HomeAssistant at the various services, as it has built in support for all of these things. 1. https://www.adafruit.com/product/5835 https://www.adafruit.com/product/5835 2. https://esphome.io/ https://esphome.io/ 3. https://gist.github.com/tedivm/2217cead94cb41edb2b50792a8bea8e6 https://gist.github.com/tedivm/2217cead94cb41edb2b50792a8bea... 4. https://github.com/BigBobbas/ESP32-S3-Box3-Custom-ESPHome/ https://github.com/BigBobbas/ESP32-S3-Box3-Custom-ESPHome/ 5. https://github.com/rhasspy/wyoming-faster-whisper https://github.com/rhasspy/wyoming-faster-whisper 6. https://github.com/rhasspy/wyoming-piper https://github.com/rhasspy/wyoming-piper 7. https://ollama.com/ https://ollama.com/
- thebruce87m 1y ago> 1. https://www.adafruit.com/product/5835 https://www.adafruit.com/product/5835 The nails in the video made me laugh
- mrandish 1y agoI run Home Assistant on an RPi4 and have an ESP32-based Core2 with mic (https://shop.m5stack.com/products/m5stack-core2-esp32-iot-development-kit https://shop.m5stack.com/products/m5stack-core2-esp32-iot-de...), along with a 16GB 4070 Ti Super in an always-on Windows system I only use for occasional gaming and serving media. I'd love to set up something like you have. Can you recommended a starting place, or ideally, a step-by-step tutorial? I've never set up any AI system. Would you say setting up such a self-hosted AI is at a point now where an AI novice can get an AI system installed and integrated with an existing Home Assistant install in a couple hours?
- Implicated 1y agoI mean - the AI itself will help you get all that setup. Claude code is your friend. I run proxmox on an old Dell R710 in my closet that hosts my homeassistant (amongst others) VM and then I've setup my "gaming" PC (which hasn't done any gaming in quite some time) to dual boot (Windows or Deb/Proxmox) and just keep it booted into Deb as another proxmox node. That PC also has a 4070 Super that I have setup to passthru to a VM and on that VM I've got various services utilizing the GPU. This includes some that are utilized by my hetzner bare metal servers for things like image/text embeddings as well as local LLM use (though, rather minimal due to VRAM constraints) and some image/video object detection stuff with my security cameras (slowly working on a remote water gun turret to keep the racoons from trying to eat the kittens that stray cats keep having in my driveway/workshop). Install claude code (or, opencode, it's also good) - use Opus (get the max plan) and give it a directory that it can use as it's working directory (don't open it in ~/Documents and just start doing things) and prompt it with something as simple as this: "I have an existing home assistant setup at home and I'd like to determine what sort of self-hosted AI I could setup and integrate with that home assistant install - can you help me get started? Please also maintain some notes in .md files in this working directory with those note files named and organized as you see appropriate so that we can share relevant context and information with future sessions. (example: Hardware information, local urls, network layout, etc) If you're unsure of something, ask me questions. Do not perform any destructive actions without first confirming with me." Plan mode. _ALWAYS_ use plan mode to get the task setup, if there's something about the plan you don't like, say no and give it notes - it will return with a new plan. Eventually agree to the plan when it's right - then work through that plan not in plan mode, but if it gets off the plan, get back in plan mode to get the/a plan set and then again let it go and just steer it in regular mode.
- state_less 1y agoThat's great to hear. I was mostly impressed with Qwen3 coder on my 4090, but am hobbled by the small memory footprint of the single card. What motherboard are you using with your 3090s? Like the others, I too am curious about those esp32s and what software you run on them. Keep up the good hacking - it's been fun to play with this stuff!
- tedivm 1y agoI actually am not using the 3090s as one unit. I have Qwen3-30B-A3B as my primary model and it fits on a single GPU, then I have all the TTS/STT on the other GPU.
- crorella 1y agoomg, this is something I've had in mind for quite some time, I even bought some i2s devices to test it out. Do you have some pointers on how to do it?
- bityard 1y agoCan you tell me about these voice satellites?
- mudkipdev 1y agoDo you also add custom tools to turn on/off the lights?
- bilbo0s 1y agoAmericans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it Wouldn't worry about that, I'm pretty sure the government is going to ban running Chinese tech in this space sooner or later. And we won't even be able to download it. Not saying any of the bans will make any kind of sense, but I'm pretty sure they're gonna say this is a "strategic" space. And everything else will follow from there. Download Chinese models while you can.
- Sanzig 1y agoWhen DeepSeek first hit the news, an American senator proposed adding it to ITAR so they could send people to prison for using it. Didn't pass, thankfully.
- thoroughburro 1y agoIf it does in the future, do we just hope it won’t be retroactive? Is this water boiling yet?
- dingnuts 1y agoex post facto law is explicitly banned in the US Constitution
- CamperBob2 1y agoDogs can't play basketball, either, but we've sure been getting dunked on a lot lately.
- jondwillis 1y agoHistory is littered with unconstitutional, enforced laws, as well. Watched a lot of Ken Burns docs this weekend while sick. “The West” has quite a few examples.
- Light_Hope 1y ago
- deleted 1y ago[deleted]
- nerdsniper 1y agoWhen has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).
- state_less 1y agoTo be fair, it isn't $1000-2000 extra, it's the new laptop/pc you just bought that is powerful enough (now, or in the near future) to run these open weight models.
- echelon 1y agoEase of use is a major issue. What percentage of the people that you know are able to install python and dependencies plus the correct open weights models? I'd wager most of your parents can't do it. Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.
- wiredpancake 1y agoThat sounds like all software. We can make better software.
- nerdsniper 1y agoWiredpancake got flagged to death but they’re right. MacWhisper provides a great example of good value for dead-simple user-friendly on-device processing.
- nickthegreek 1y agoinstalling LM Studio is easy and it walks you through choosing a model as well. It is actually well within many people's abilities.
- MangoToupe 1y agoYou mean like a home with a yard large enough to keep the neighbors out of sight? Granted, based on how annoyingly chill we are with advertisements and government surveillance, I suppose this desire for privacy never extended beyond the neighbors.
- deleted 1y ago[deleted]
- moffkalast 1y agoThere is some irony about buying Chinese hardware to run American software on it for the past decade(s), and now the exact reverse.
- fragmede 1y agoIs that irony? Regardless, it's hilarious. All of the upvotes to you!
- Aurornis 1y ago> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscription is a bargain and they automatically get access to new versions as they come. I think the at-home self hosting hobby is interesting, but it’s never going to be a mainstream thing.
- rafterydj 1y agoI agree the market is niche atm, but I can't help but disagree with your outlook long term. Self hosted models don't have the problems ChatGPT subscribers are facing with models seemingly performing worse over time, they don't need to worry about usage quotas, they don't need to worry about getting locked out of their services, etc. All of these things have a dark side, though; but it's likely unnecessary for me to elaborate on that.
- reilly3000 1y agoThere is going to be a big market for private AI appliances, in my estimation at least. Case in point: I give Gmail OAuth access to nobody. I nearly got burned once and I really don’t want my entire domain nuked. But I want to be able to have an LLM do things only LLMs can do with my email. “Find all emails with ‘autopay’ in the subject from my utility company for the past 12 months, then compare it to the prior year’s data.” GPT-OSS-20b tried its best but got the math obviously wrong. Qwen happily made the tool calls and spat out an accurate report, and even offered to make a CSV for me. Surely if you can’t trust npm packages or MS to not hand out god tokens to any who asks nicely, you shouldn’t trust a random MCP server with your credentials or your model. So I had Kilocode build my own. For that use case, local models just don’t quite cut it. I loaded $10 into OpenRouter, told it what I wanted, and selected GPT5 because it’s half off this week. 45 minutes, $0.78, and a few manual interventions later I had a working Gmail MCP that is my very own. It gave me some great instructions on how to configure an OAuth app in GCP, and I was able to get it running queries within minutes from my local models. There is a consumer play for a ~$2499-$5000 box that can run your personal staff of agents on the horizon. We need about one more generation of models and another generation of low-mid inference hardware to make it commercially feasible to turn a profit. It would need to pay for itself easily in the lives of its adopters. Then the mass market could open up. A more obvious path goes through SMBs who care about control and data sovereignty. If you’re curious, my power bill is up YoY, but there was a rate hike, definitely not my 4090;).
- mbac32768 1y agositting here in the US, reading that China is strongly urging the adoption of Linux and pushing for open CPU architectures like RISC-V and also self-hosted open models are we the baddies??
- BobbyJo 1y agoIf there is a walled garden, and you aren't in it, you'll probably push for the walls to come down. No moral basis needed.
- vachina 1y agosilent majority of HN
- h4ny 1y agoCould you elaborate on what you mean by "moral basis" in your comment?
- BobbyJo 1y agoI mean China's push for open weights/source/architecture probably has more to do with them wanting legal access to markets than it does with those things being morally superior.
- wiz21c 1y agoIf by being selfish they end up doing morally superior thing, then, I much prefer to go with the Chinese. Even more so now that Trump is in command.
- amunozo 1y agoOf course, but that translates in a benefit for most people, even for Americans. In my case (European), I cannot but support the Chinese companies in this respect, as we would be especially in trouble if the common models are the norm.
- 1y ago
- deleted 1y ago[deleted]
- logicallee 1y ago>Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster So did you run the model offline on your own computer and get realtime audio? Can you tell me the GPU or specifications you used? I inquired with ChatGPT: https://chatgpt.com/share/68d23c2c-2928-800b-bdde-040d8cb40b93 https://chatgpt.com/share/68d23c2c-2928-800b-bdde-040d8cb40b... It seems it needs around a $2,500 GPU, do you have one? I tried Qwen online via its website interface a few months ago, and found it to be very good. I've run some offline models including Deepseek-R1 70B on CPU (pretty slow, my server has 128 GB of RAM but no GPU) and I'm looking into what kind of setup I would need to run an offline model on GPU myself.
- threeducks 1y ago> So did you run the model offline on your own computer and get realtime audio? At the top of the README of the GitHub repository, there are a few links to demos where you can try the model. > It seems it needs around a $2,500 GPU You can get a used RTX 3090 for about $700, which has the same amount of VRAM as the RTX 4090 in your ChatGPT response. But as far as I can tell, quantized inference implementations for this model do not exist yet.
- powerapple 1y agoIs there a AI market for open weights? Companies like Alibaba, Tencent, Meta or Microsoft makes a lot sense. They can build on open weights, and not losing values, potentially beneficial for share prices. The only winner is application and cloud providers, I don't see how they can make money from the weights itself to be honest.
- a2128 1y agoIt promotes an open research environment where external researchers have the opportunity to learn, improve and build. And it keeps the big companies in check, they can't become monopolies or duopolies and increase API prices (as is usually the playbook) if you can get the same quality responses from a smaller provider on OpenRouter
- amunozo 1y agoI don't know if there is a market for it, but I know that open weights puts pressure on the closed-models companies into releasing their weights and losing their privileged situations.
- Cheer2171 1y agoThe only money to be made is in compute, not open weights themselves. What point is a market when a commons like huggingface or modelscope? Alibaba made modelscope to compete with HF, and that's a commons not a market either, if that tells you anything. By analogy, you can legally charge for copies of your custom Linux distribution, but what's the point when all the others are free?
- pg3uk 1y agoThe US is probably ahead but they're so obsessed with moats, IP and safety that their lagginess is self imposed. China has nothing to lose and everything to gain by releasing stuff openly. Once China figures put how to make high performance FPGA chips really cheap, its game over for the US. The only power the US has is over GPU supply...and even then its pretty weak. Not to mention NVIDIA crippling its own country with low VRAM cards. China is taking older cards, stripping the RAM and upgrading other older cards.