5 ms·
So weird/cool/interesting/cyberpunk that we have stuff like this in the year of our Lord 2026: ├── MEMORY.md # Long-term knowledge (auto-loaded e
by dvt 8mo ago
So weird/cool/interesting/cyberpunk that we have stuff like this in the year of our Lord 2026:
├── MEMORY.md # Long-term knowledge (auto-loaded each session)
├── HEARTBEAT.md # Autonomous task queue
├── SOUL.md # Personality and behavioral guidance
Say what you will, but AI really does feel like living in the future. As far as the project is concerned, pretty neat, but I'm not really sure about calling it "local-first" as it's still reliant on an `ANTHROPIC_API_KEY`.
I do think that local-first will end up being the future long-term though. I built something similar last year (unreleased) also in Rust, but it was also running the model locally (you can see how slow/fast it is here[1], keeping in mind I have a 3080Ti and was running Mistral-Instruct).
I need to re-visit this project and release it, but building in the context of the OS is pretty mindblowing, so kudos to you. I think that the paradigm of how we interact with our devices will fundamentally shift in the next 5-10 years.
[1] https://www.youtube.com/watch?v=tRrKQl0kzvQ https://www.youtube.com/watch?v=tRrKQl0kzvQ
- atmanactive 8mo ago> but I'm not really sure about calling it "local-first" as it's still reliant on an `ANTHROPIC_API_KEY`. See here: https://github.com/localgpt-app/localgpt/blob/main/src%2Fagent%2Fproviders.rs#L222 https://github.com/localgpt-app/localgpt/blob/main/src%2Fage...
- nodesocket 8mo agoWhat reasonable comparable model can be run locally on say 16GB of video memory compared to Opus 4.6? As far as I know Kimi (while good) needs serious GPUs GTX 6000 Ada minimum. More likely H100 or H200.
- lodovic 8mo agoI made something similar to this project, and tested it against a few 3B and 8B models (Qwen and Ministral, both the instruction and the reasoning variants). I was pleasantly surprised by how fast and accurate these small models have become. I can ask it things like "check out this repo and build it", and with a Ralph strategy eventually it will succeed, despite the small context size.
- mixermachine 8mo agoNothing will come close to Opus 4.6 here. You will be able to fit a destilled 20B to 30B model on your GPU. Gpt-oss-20B is quite good in my testing locally on a Macbook Pro M2 Pro 32GB. The bigger downside, when you compare it to Opus or any other hosted model, is the limited context. You might be able to achieve around 30k. Hosted models often have 128k or more. Opus 4.6 has 200k as its standard and 1M in api beta mode.
- zozbot234 8mo agoThere are local models with larger context, but the memory requirements explode pretty quickly so you need to lower parameter count or resort to heavy quantization. Some local inference platforms allow you to place the KV cache in system memory (while still otherwise using GPU). Then you can just use swap to allow for even very long contexts, but this slows inference down quite a bit. (The write load on KV cache is just appending a KV vector per inferred token, so it's quite compatible with swap. You won't be wearing out the underlying storage all that much.)
- PeterStuer 8mo agoNothing close to Opus is available in open weights. That said, do all your tasks need the power of Opus?
- lxgr 8mo agoThe problem is that having to actively decide when to use Opus defeats much of the purpose. You could try letting a model decide, but given my experience with at least OpenAI’s “auto” model router, I’d rather not.
- PeterStuer 8mo agoI also don't like having to think about it, and if it were free, I would not bother even though keeping up a decent local alternative is a good defensive move regardless. But let's face it. For most people Opus comes at a significant financial cost per token if used more than very casual, so using it for rather trivial or iterative tasks that nevertheless consume a lot of those is something to avoid.
- berkes 8mo agoDevstral¹ has very good models that can be run locally. They are in the top of open models, and surpass some closed models. I've been using devstral, codestral and Le Chat exclusively for three months now. All from misteals hosted versions. Agentic, as completion and for day-to-day stuff. It's not perfect, but neither is any other model or product, so good enough for me. Less anecdotal are the various benchmarks that put them surprisingly high in the rankings ¹https://mistral.ai/news/devstral https://mistral.ai/news/devstral
- halJordan 8mo agoYou absolutely do not have to use a third party llm. You can point it to any openai/anthropic compatible endpoint. It can even be on localhost.
- dvt 8mo agoAh true, missed that! Still a bit cumbersome & lazy imo, I'm a fan of just shipping with that capability out-of-the-box (Huggingface's Candle is fantastic for downloading/syncing/running models locally).
- embedding-shape 8mo agoAh come on, lazy? As long as it works with the runtime you wanna use, instead of hardcoding their own solution, should work fine. If you want to use Candle and have to implement new architectures with it to be able to use it, you still can, just expose it over HTTP.
- dvt 8mo agoI think one of the major problems with the current incarnation of AI solutions is that they're extremely brittle and hacked-together. It's a fun exciting time, especially for us technical people, but normies just want stuff to "work." Even copy-pasting an API key is probably too much of a hurdle for regular folks, let alone running a local ollama server in a Docker container.
- Sharlin 8mo agoUnlike in image/video gen, at least with LLMs the "best" solution available isn’t a graph/node-based interface with an ecosystem of hundreds of hacky undocumented custom nodes that break every few days and way too complex workflows made up of a spaghetti of two dozen nodes with numerous parameters each, half of which have no discernible effect on output quality and tweaking the rest is entirely trial and error.
- 8mo ago
- IhateAI 8mo agolocal first is not the future, lmfao, maybe in 10-20 years. It currently cost ~80k-100k to run a pretty meh Kimi 2.5 at decent tok p/s, which is rather useless anyways. And that doesn't allow you to run any multi-agent sessions. By time hardware costs shrink to allow you to run useful models, concurrently in multi agent environments, they'll have already devalued labor on a scale never before seen.The layoffs and labor will cause us all to work for morsels, on whatever work opportunities remain. Eventually you'll beg to fight in a war. LLMs are only here to attack labor, devalue the working class and eventually make us useless to the ruling class. LLMs do not create opportunities/jobs, they replace the inputs to labor, humans. That's their only purpose. But I guess most llm-kiddies think they're going to vibe code their way out of the working class with Anthropic's latest slop offering. Good luck with that. In 5 years your labor will be worth a 1/4 maybe 1/2 of what it is now, and that vibe coded startup of yours will have been made 5000x times over by every other delusional llm-kiddie. Have fun with your GPU, you won't be able to afford a 60 series, if they even make one, and it certainly won't be powerful enough to pull you out of the black mirror episode we're heading towards. I recommend learning and not frying your brain with "Think for me Saas", and not being dependent on Meta or Alibaba open sourcing some model that allows you to compete with them.
- fy20 8mo ago> Say what you will, but AI really does feel like living in the future. Love or hate it, the amount of money being put into AI really is our generation's equivalent of the Apollo program. Over the next few years there are over 100 gigawatt scale data centres planned to come online. At least it's a better use than money going into the military industry.
- pwndByDeath 8mo agoLoL, don't worry they are getting their dose of the snakeoil too
- jazzyjackson 8mo agoWhat makes you think AI investment isn't a proxy for military advantage? Did you miss the saber rattling of anti-regulation lobbying, that we cannot pause or blink or apply rules to the AI industry because then China would overtake us?
- T-A 8mo agoThe Apollo program was peanuts in comparison: https://www.wsj.com/tech/ai/ai-spending-tech-companies-compared-02b90046 https://www.wsj.com/tech/ai/ai-spending-tech-companies-compa... https://www.reuters.com/graphics/USA-ECONOMY/AI-INVESTMENT/gkvlqbgxkpb/ https://www.reuters.com/graphics/USA-ECONOMY/AI-INVESTMENT/g...
- adammarples 8mo agoYou know they will never come on line. A lot of it is letters of intention to invest with nothing promised, mostly to juice the circular share price circuils.
- ryan_n 8mo agoMost of these AI companies are part of the military industry. So the money is still going there at the end of the day.
- jazzyjackson 8mo agoIMHO it doesn't make sense, financially and resource wise to run local, given the 5 figure upfront costs to get an LLM running slower than I can get for 20 USD/m. If I'm running a business and have some number of employees to make use of it, and confidentiality is worth something, sure, but am I really going to rely on anything less then the frontier models for automating critical tasks? Or roll my own on prem IT to support it when Amazon Bedrock will do it for me?
- zozbot234 8mo agoIt starts making a lot of sense if you can run the AI workloads overnight on leaner infrastructure rather than insist on real-time response.
- Sharlin 8mo agoThat’s probably true only as long as subscription prices are kept artificially low. Once the $20 becomes $200 (or the fast-mode inference quotas for cheap subs become unusably small), the equation may change.
- berkes 8mo agoThis field is highly competitive. Much more than I expected it to. I thought the barrier to entry was so high, only big tech could seriously join the race, because of costs, or training data etc. But there's fierce competition by new or small players (deepseek, Mistral etc), many even open source. And Icm convinced they'll keep the prices low. A company like openai can only increase subscriptions x10 when they've locked in enough clients, have a monopoly or oligopoly, or their switching costs are multitudes of that. So currently the irony seems to be that the larger the AI company, the more loss they're running at. Size seems to have a negative impact on business. But the smaller operators also prevent companies from raising prices to levels at which they make money.
- Sharlin 8mo agoThere's no way around the cost of electricity, at least in the short term. Nobody has come up with a way to meaningfully scale capacity without scaling parameter count (≈energy use). Everybody seems to agree that the newest Claudes are the only coding models capable of some actually semi-challenging tasks, and even those are prone to all the usual failure modes and require huge amounts of handholding. No smaller models seem to get even close.
- __mharrison__ 8mo agoI'm playing with local first openclaw and qwen3 coder next running on my LAN. Just starting out but it looks promising.
- bluerooibos 8mo agoOn what sort of hardware/RAM? I've been trying ollama and opencode with various local models on a 16Gb RAM, but the speed, and accuracy/behaviour just isn't good enough yet.
- __mharrison__ 8mo agoDGX Spark (128gb)
- backscratches 8mo agoYes this is not local first, the name is bad.
- lxgr 8mo agoTo be precise, it’s exactly as local first as OpenClaw (i.e. probably not unless you have an unusually powerful GPU).
- backscratches 8mo agoYes but OpenClaw (which is a terrible name for other reasons) doesn't have "local" in the name and so is not misleading.
- outofpaper 8mo agoAs misleading. Lots of their marketing push or at least thr ClawBros pitch it as running local on your MacMini.
- lxgr 8mo agoTo be fair, you do keep significantly more control of your own data from a data portability perspective! A MEMORY.md file presents almost zero lock-in compared to some SaaS offering. Privacy-wise, of course, the inference provider sees everything.
- jagged-chisel 8mo agoTo be clear: keeping a local copy of some data provides not control over how the remote system treats that data once it’s sent.
- lxgr 8mo agoWhich is what I said in my second sentence.
- croes 8mo ago> but AI really does feel like living in the future. Got the same feeling when I put on the Hololens for the first time but look what we have now.
- mycall 8mo agoWhat does ANTHROPIC bring to this project that a local LLM cannot, e.g. Gwen3 Coder Next?