8 ms·
Nanobot: Ultra-Lightweight Alternative to OpenClaw
- johaugum 8mo agoSkimmed the repo, this is basically the irreducible core of an agent: small loop, provider abstraction, tool dispatch, and chat gateways . The LOC reduction (99%, from 400k to 4k) mostly comes from leaving out RAG pipelines, planners, multi-agent orchestration, UIs, and production ops.
- baby 8mo agoRAG seems odd when you can just have a coding agent manage memory by managing folders. Multi agent also feels weird when you have subagents.
- antirez 8mo agoTotally useless indeed.
- rando77 8mo agoI've been leaning towards multi agent because sub agent relies on the main agent having all the power and using it responsibly.
- andai 8mo agoWhat does that mean?
- rando77 8mo agoHave you come across the lethal trifecta [1]? I'm interested in decomposing tasks to avoid any agent doing all three and having limited data pipes between them. [1] https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
- PlatoIsADisease 8mo agoInteresting. I guess RAG is faster? But I'm realizing I'm outdated now.
- lxgr 8mo agoNo, RAG is definitely preferable once your memory size grows above a few hundred lines of text (which you can just dump into the context for most current models), since you're no longer fighting context limits and needle-in-a-haystack LLM retrieval performance problems.
- Aurornis 8mo ago> once your memory size grows above a few hundred lines of text (which you can just dump into the context for most current models) A few hundred lines of text is nothing for current LLMs. You can dump the entire contents of The Great Gatsby into any of the frontier LLMs and it’s only around 70K tokens. This is less than 1/3 of common context window sizes. That’s even true for models I run locally on modest hardware now. The days of chunking everything into paragraphs or pages and building complex workflows to store embeddings, search, and rerank in a big complex pipeline are going away for many common use cases. Having LLMs use simpler tools like grep based on an array of similar search terms and then evaluating what comes up is faster in many cases and doesn’t require elaborate pipelines built around specific context lengths.
- deleted 8mo ago[deleted]
- lxgr 8mo agoYes, but how good will the recall performance be? Just because your prompt fits into context doesn't mean that the model won't be overwhelmed by it. When I last tried this with some Gemini models, they couldn't reliably identify specific scenes in a 50K word novel unless I trimmed down the context to a few thousands of words. > Having LLMs use simpler tools like grep based on an array of similar search terms and then evaluating what comes up is faster in many cases Sure, but then you're dependent on (you or the model) picking the right phrases to search for. With embeddings, you get much better search performance.
- simonw 8mo agoYeah, vector embeddings based RAG has fallen out of fashion somewhat. It was great when LLMs had 4,000 or 8,000 token context windows and the biggest challenge was efficiently figuring out the most likely chunks of text to feed into that window to answer a question. These days LLMS all have 100,000+ context windows, which means you don't have to be nearly as selective. They're also exceptionally good at running search tools - give them grep or rg or even `select * from t where body like ...` and they'll almost certainly be able to find the information they need after a few loops. Vector embeddings give you fuzzy search, so "dog" also matches "puppy" - but a good LLM with a search tool will search for "dog" and then try a second search for "puppy" if the first one doesn't return the results it needs.
- y1n0 8mo agoContext rot is still a problem though, so maybe vector search will stick around in some form. Perhaps we will end up with a tool called `vector grep` or `vg` that handles the vectorized search independent of the agent.
- visarga 8mo agoThe fundamental problem wit RAG is that it extracts only surface level features, "31+24" won't embed close to "55", while "not happy" will be close to "happy". Another issue is that embedding similarity does not indicate logical dependency, you won't retrieve the callers of a function with RAG, you need a LLM or code for that. Third issue is chunking, to embed you need to chunk, but if you chunk you exclude information that might be essential. The best way to search I think is a coding agent with grep and file system access, and that is because the agent can adapt and explore instead of one shotting it. I am making my own search tool based on the principle of LoD (level of detail) - any large text input can be trimmed down to about 10KB size by doing clever trimming, for example you could trim the middle of a paragraph keeping the start and end, or you could trim the middle of a large file. Then an agent can zoom in and out of a large file. It skims structure first, then drills into the relevant sections. Using it for analyzing logs, repos, zip files, long PDFs, and coding agent sessions which can run into MB size. Depending on content type we can do different types of compression for code and tree structured data. There is also a "tall narrow cut" (like cut -c -50 on a file). The promise is - any size input fit into 10KB "glances" and the model can find things more efficiently this way without loading the whole thing.
- m00dy 8mo agoRAG is broken when you have too much data.
- thunky 8mo agoGemini with Google search is RAG using all public data, and it isn't broken.
- fhd2 8mo agoIt's not tool use with natural language search queries? That's what I'd expect.
- kaicianflone 8mo agoIt is tool use with natural language search queries but going down a layer they are searched on a vector DB, very similar to RAG. Essentially Google RankBrain is the very far ancestor to RAG before compute and scaling.
- thunky 8mo agoIt's RAG via tool use, where the storage and retreival method is an implementation detail. I'm not a huge fan of the term RAG though because if you squint almost all tool use could be considered RAG. But if you stick with RAG being a form of "knowledge search" then I think Google search easily fits.
- PlatoIsADisease 8mo agoCant you make thresholds higher? Hmm... I guess not, you might want all that data. Super interesting topic. Learning a lot.
- plingamp 8mo agoSpecifically when the document number reaches around 10k+, a phenomenon called "Semantic Collapse" occurs. https://dho.stanford.edu/wp-content/uploads/Legal_RAG_Hallucinations.pdf https://dho.stanford.edu/wp-content/uploads/Legal_RAG_Halluc...
- naasking 8mo agoUnless I'm misunderstanding what they are, planners seem kind of important.
- johaugum 8mo agoAs you mentioned, that depends on what you mean by planners. An LLM will implicitly decompose a prompt into tasks and then sequentially execute them, calling the appropriate tools. The architecture diagram helpfully visualizes this [0] Here though, planners means autonomous planners that exist as higher level infrastructure, that does external task decomposition, persistent state, tool scheduling, error recovery/replanning, and branching/search. Think a task like “Prompt: “Scan repo for auth bugs, run tests, open PR with fixes, notify Slack.” that just runs continuously 24/7, that would be beyond what nanobot could do. However, something like “find all the receipts in my emails for this year, then zip and email them to my accountant for my tax return” is something nanobot would do. [0] https://github.com/HKUDS/nanobot/blob/main/nanobot_arch.png https://github.com/HKUDS/nanobot/blob/main/nanobot_arch.png
- naasking 8mo agoSure, instruction tuned models implicitly plan, but they can easily lose the plot on long contexts. If you're going to have an agent running continuously and accumulating memory (parsing results from tool use, web fetches, previous history, etc.), then plan decomposition, persistence and error recovery seems like a good idea, so you can start subagents with fresh contexts for task items and they stay on task or can recover without starting everything over again. Also seems better for cost since input and output contexts are more bounded.
- skybrian 8mo agoI don’t know what these planners do, but I’ve had reasonably good luck asking a coding agent to write a design doc and then reviewing it a few times.
- jannniii 8mo agoOkay so is this ”inspired” by nanoclaw that was featured here two days ago?
- jimmcslim 8mo agoHah, I was looking at this and going "wasn't this up on HN front page just a few days ago?" and has completely missed Nanoclaw vs Nanobot.
- reustle 8mo agoFor the curious https://github.com/gavrielc/nanoclaw https://github.com/gavrielc/nanoclaw
- Copenjin 8mo agoIt's so simple conceptually that I'm not sure if you even need to be inspired by something other than the original cladwbot.
- vanillameow 8mo agoYeah I mean idk, my takeaway from OpenClaw was pretty much the same - why use someone's insane vibecoded 400k LoC CLI wrapper with 50k lines of "docs" (AI slop; and another 50k Chinese translation of the same AI slop) when I can just Claude Code myself a custom wrapper in 30 mins that has exactly what I need and won't take 4 seconds to respond to a CLI call. But my reaction to this project is again: Why would I use this instead of "vibecoding" it myself. It won't have exactly what I need, and the cost to create my own version is measured in minutes. I suspect many people will slowly come to understand this intrinsic nature of "vibecoded software" soon - the only valuable one is one you've made yourself, to solve your own problems. They are not products and never will be.
- pelagicAustral 8mo agoI mean, in not vibecoding it yourself you are already saving tokens... Personally, I see no benefit in having an instance of something like this... so, I wouldn't spend tokens, and I wouldn't spend server-time, or any other resource into it, but a lot of people seem to have found a really nice alternative to actually having to use their brains during the day.
- vanillameow 8mo agoI do see the potential in something like OpenClaw, personally, but more as a kind of interface for a collection of small isolated automations that _could_ be loosely connected via some type of memory bank (whether that's a RAG or just text files or a database or whatever). Not all of these will require LLMs and certainly none of them will require vibecoding at all if you have infinite time; But the reality is I don't have infinite time, and if I have 300 small ideas and I can only implement my like 10 of them a week by myself, I'd personally rather automate 30 more than just not have them at all, you know? But I am talking about shell scripts here, cronjobs, maybe small background services. And I would never dare publish these as public applications or products. Both because I feel no pride about having "made" these - because, you know, I haven't, the AI did - and because they just aren't public facing interfaces. I think the main issue at the moment is that so many devs are pretending that these vibecoded projects are "products". They are not. They are tailor-made, non-recyclable throwaway software for one person: The creator. I just see no world at the moment where I have any plausible reason to use someone else's vibecoded software.
- loveparade 8mo agoWhat are people using these things for? The use cases I've seen look a bit contrived and I could ask Claude or ChatGPT to do it directly
- dominicq 8mo agoYeah, I don't get it either. Deploy a VM that runs an LLM so that I can talk to it via Telegram... I could just talk to it through an app or a web interface. I'm not even trying to be snarky, like what the hell even is the use case?
- BoredPositron 8mo agoIt's not even an LLM it's just to pipe api calls.
- xylo 8mo agoDifference is that openclaw is not LLM but engine that spawns up agent that interact with LLM and the system its installed on. It can have full access to the system it’s running on. So it can browse internet via browser, run cli commands, api’s via skills etc. Idea is to act like a Jarvis personal assistant. You tell what to do via chat e.g telegram, then it does it for you.
- gergo_b 8mo agoI have no idea. the single thing I can think of is that it can have a memory.. but you can do that with even less code. Just get a VPS. create a folder and run CC in it, tell it to save things into MD files. You can access it via your phone using termux.
- deleted 8mo ago[deleted]
- sReinwald 8mo agoYou could, but Claude Code's memory system works well for specialized tasks like coding - not so much for a general-purpose assistant. It stores everything in flat markdown files, which means you're pulling in the full file regardless of relevance. That costs tokens and dilutes the context the model actually needs. An embedding-based memory system (letta, mem0, or a self-built PostgreSQL + pgvector setup) lets you retrieve selectively and only grab what's relevant to the current query. Much better fit for anything beyond a narrow use case. Your assistant doesn't need to know your location and address when you're asking it to look up whether sharks are indeed older than trees, but it probably should know where you live when you ask it about the weather, or good Thai restaurants near you.
- FergusArgyll 8mo agoThe main novelty I see in openclaw is the amount of channels and how easy it is to set them up. This just has whatsapp, telegram & feishu
- tunney 8mo agoHas anyone managed to get the WhatsApp integration working and chatting that way?
- Aeroi 8mo agocan anyone breakdown a comparison of multi-agent vs subagent? looking for pro's and cons.
- Tepix 8mo agoWhat are your solutions for if your AI bot wants to leak your credentials?
- manwithmanyface 8mo agoIs this something I run for my company in Slack, where employees send messages and the LLM processes the text, uses the functions I created to handle different tasks, and then responds back?
- sally-suite 8mo agoNot bad, but I’m a bit skeptical. Is it mainly about the way of working in IM?
- yberreby 8mo agoWatching the OpenClaw/Molbot craze has been entertaining. I wouldn't use it - too much code, changing too quickly, with too little regard for security - but it has inspired me. I often have ideas while cleaning around, cooking, etc. Claude Code (with Opus 4.5) is very capable. I've long wanted to get Claude Code working hands-free. So I took an afternoon and rolled my own STT-TTS voice stack for Claude Code. The voice stack runs locally on my M4 Pro and is extremely fast. For Speech to Text, Parakeet v3 TDT: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 For Text to Speech, Pocket TTS: https://github.com/kyutai-labs/pocket-tts https://github.com/kyutai-labs/pocket-tts Custom MCP to hook this into Claude Code, with a little bit of hacking around to get my AirPods' stem click to be captured. I'm having Claude narrate its thought process and everything it's doing in short, frequent messages, and I can interrupt it at any time with a stem click, which starts listening to me and sends the message once a sufficiently long pause is detected. I stream the Claude Code session via AirPlay to my living room TV, so that I don't have to get close to the laptop if I need extra details about what it's doing. Yesterday, I had it debug a custom WhatsApp integration (via [1]) hands-free while brushing my teeth. It can use `osascript` for OS integration, browse the web via Claude Code's builtin tools... My back is thankful. This is really fun. [1]: https://github.com/jlucaso1/whatsapp-rust https://github.com/jlucaso1/whatsapp-rust
- gdhkgdhkvff 8mo agoOn one hand, I think this project is super cool and something I would use and/or would have loved to build myself for my own use. On the other hand, it makes me wonder if we’re just heading for a future where everyone is just always working, at all times, even while doing other things. “Wow look at our daughter taking her first steps! She’s doing so… wait hold on… No, Claude. I said to name the class “potatoes”, not “‘pot’ followed by eight ‘O’s,” you dumb robot!”
- volkk 8mo agowe kind of already are with our phones and Slack, the difference at this point is negligible. i personally won't have airpods in 24/7 with my kid (or ever) so if i were doing something like this, it would be through my phone, which is already something i use fairly often. not too much difference there IMO (at least anecdotally speaking)
- lxgr 8mo agoCan this be sandboxed? I've been running OpenClaw in a VM on macOS, which seems more resource intensive than necessary.
- cpursley 8mo agoI'd like to see one of these in Rust (over Python, Node, etc) and in Apple's container environment.
- raphaelmolly8 8mo ago[dead]
- mrklol 8mo agoDidn’t openclaw switch to vector based because it used way less tokens as it always loaded all memories? Seems way more efficient
- pawelduda 8mo agoWhat? OpenClaw has 450kLoC? Why?
- halfax 8mo agoBottom Line HAL‑AI‑2 is a real system. Nanobot is a toy. They are not peers. They are not even in the same category. Nanobot is useful only as a conceptual sketch of an agent loop. HAL‑AI‑2 is the substrate you’ve been building toward for months.
- resonious 8mo agoI hate to side-track like this but I'm having trouble understanding the architecture diagram. LLM has 2 arrows to Tools - what does that mean? Similarly, Tools has both a doublesided arrow and an outgoing arrow to Context. Chat Apps having outgoing arrows to both Message and LLM also kinda tripped me up but I suppose you could say it's because the apps both provide messaging and context for the LLM.
- fathermarz 8mo agoI have been inspired by all the use cases that are popping up from a proactive assistant, but lightweight is the last thing I would want when it comes to locking it down. I started building my own version and before I even think about letting it loose, every facet needs to be designed and thought out. I have more tests than these lightweight libraries have code. To me I don’t care about the size, I care about not getting wrecked.