7 ms·
Don't trust AI agents
- sarkarsh 7mo ago[dead]
- formerly_proven 7mo agod'uh
- smallpipe 7mo agoDocker is not a security boundary. You’re one prompt injection away from handing over your gmail cookie.
- benatkin 7mo agoNo, but Podman is. The recent escapes at the actual container level have been pretty edge case. It's been some years since a general container escape has been found. Docker's CVE-2025-9074 was totally unnecessary and due to Docker being Docker.
- eyberg 7mo agoNo they have not been. There were at least 16 container escapes last year - at least 8 of them were at the runtime layer. I personally spent way too much time looking at this in the past month: https://nanovms.com/blog/last-year-in-container-security https://nanovms.com/blog/last-year-in-container-security runc: https://www.cve.org/CVERecord?id=CVE-2025-31133 https://www.cve.org/CVERecord?id=CVE-2025-31133 nvidia: https://www.cve.org/CVERecord?id=CVE-2025-23266 https://www.cve.org/CVERecord?id=CVE-2025-23266 runc: https://www.cve.org/CVERecord?id=CVE-2025-52565 https://www.cve.org/CVERecord?id=CVE-2025-52565 youki: https://www.cve.org/CVERecord?id=CVE-2025-54867 https://www.cve.org/CVERecord?id=CVE-2025-54867 Also, last time I checked podman uses runc by default.
- jrpear 7mo agoIt looks to me like what is called a "container escape" in this context isn't necessarily as bad as it seems. For example, in the advisory for CVE-2025-31133 affecting runc[1]: > Container Escape: ...Thus, the attacker can simply trigger a coredump and gain complete root privileges over the host. Sounds bad. But... > this flaw effectively allows any attacker that can spawn containers (with some degree of control over what kinds of containers are being spawned) to achieve the above goals. The attacker needs already to have the capability to spawn containers! This isn't a case of "RCE within the container" -> "RCE outside the container", which is what I would think prima facie reading "container escape". I have always thought that running an untrusted image within an unprivileged container was a safe thing to do and I still believe so. [1] https://github.com/opencontainers/runc/security/advisories/GHSA-9493-h29p-rfm2 https://github.com/opencontainers/runc/security/advisories/G...
- xienze 7mo agoThe best container security in the world isn’t going to help you when the agent has credentials to third party services. Frankly, I don’t think bad actors care that much about exploiting agents to rm -rf /. It’s much more valuable to have your Google tokens or AWS credentials.
- himata4113 7mo agoMy assistant has no permissions at all and is just as useful. All it needs is todo, reminders and websearch (and maybe a browser but ymmv).
- piker 7mo ago> no permissions at all > and maybe a browser does not compute
- yyyk 7mo agoI suspect OP actually means 'cannot access anything locally' by 'no permissions'.
- himata4113 7mo agoI sometimes forget to be very clear about what I mean, too many ways to misinterpret these things.
- himata4113 7mo agoA browser doesn't magically give access to secrets, it is useful for looking up things behind a captcha.
- isodev 7mo ago> websearch (and maybe a browser Your assistant can literally be told what to do and how to hide it from you. I know security is not a word in slopware but as a high-level refresher - the web is where the threats are.
- himata4113 7mo agoWhat will it do... leak my todo...? lol. It's in a pod with zero permissions, secrets or access to the local network. It's also restarted daily incase somehow someone manages to escape a browser.
- croes 7mo ago
- VladVladikoff 7mo agoThis doesn’t really feel like enough guardrails to prevent the type of problems we’ve seen so far. For example an agent in a single container which has access to an email inbox, can still do a lot of damage if that agent goes off the rails. We agree this agent should not be trusted, yet the ideas proposed as a solution are insufficient. We need a fundamentally different approach. Also and this is just my ignorance about Claws, but if we allow an agent permission to rewrite its code to implement skills, what stops it from removing whatever guardrails exist in that codebase?
- gronky_ 7mo agoDon’t know about other claws, with NanoClaw the agent can only rewrite code that runs inside the container. You can see here that it’s only given write access to specific directories: https://github.com/qwibitai/nanoclaw/blob/8f91d3be576b830081f2a802e2f2d426b010f8f7/src/container-runner.ts#L74 https://github.com/qwibitai/nanoclaw/blob/8f91d3be576b830081...
- float4 7mo agoWouldn't you get >50% of the usefulness and 0% of the risk if you add read+draft permissions for the email connection through a proxy or oauth permissions? Then your claw can draft replies and you have to manually review+send. It's not a perfect PA that way, but could still be better than doing everything yourself for the vast majority of people who don't have a PA anyway? It feels like, just like SWEs do with AI, we should treat the claw as an enthusiastic junior: let it do stuff, but always review before you merge (or in this case: send).
- jrecyclebin 7mo agoAgent can still "forgot password" on many accounts. Or magic link.
- deleted 7mo ago[deleted]
- coffeefirst 7mo ago
- adithyassekhar 7mo agoReally good points about ai making gigantic heaps of code no human can ever review. It's almost like bureaucracy. The systems we have in governments or large corporations to do anything might seem bloated an could be simplified. But it's there to keep a lot of people employed, pacified, powers distributed in a way to prevent hostile takeovers (crazy). I think there was a cgp grey video about rulers which made the same point. Similarly AI written highly verbose code will require another AI to review or continue to maintain it, I wonder if that's something the frontier models optimize for to keep them from going out of business. Oh and I don't mind they're bashing openclaw and selling why nanoclaw is better. I miss the times when products competed with each other in the open.
- nz 7mo agoAn interesting economic fact: Karl Marx observed that if factories keep getting more efficient, eventually, they will require fewer workers because the population is not growing quickly enough to match the increasing rate of production. This, as we have seen historically, is correct: we have fewer workers per factory and fewer factories per manufactured widget. Marx also observed that this will create mass unemployment. While this is _logically_ correct, it did not really turn out that way _historically_. Most of the manufacturing labor was replaced with bureaucratic labor (so called white-collar labor) -- all of those manufacturing firms needed to grow their internal bureaucracies to manage and direct a sprawling supply-chain.
- shich 7mo agothe trust problem cuts both ways tho — users don't trust agents, but the bigger issue is agents trusting each other. once you have multi-agent pipelines, you're one rogue upstream output away from a cascade. sandboxing individual agents is table stakes; what's actually hard is defining trust boundaries between them
- medi8r 7mo agoAlso agents cannot trust any data whatsoever they add to their context. This puts reading email for example as a risk. Probably not impossible to create a worm that convinces a claw to forward it to every email address in that inbox. And then exfiltrate all the emails. Then do a bunch of password resets. Then get root access to your claw. But not just email. Github issues, wikipedia, HN etc. may be poisoned. See https://simonw.substack.com/p/the-lethal-trifecta-for-ai-agents https://simonw.substack.com/p/the-lethal-trifecta-for-ai-age... but there may be more trifectas than that in a claw driven future.
- lucrbvi 7mo agoWhy does OpenClaw have 800,000+ lines of code?? Isn't it just a connector for LLM APIs and other tools?
- zarzavat 7mo agoThey are probably counting dependencies. Also, it's vibe coded, what do you expect! I used to think that LLMs would replace humans but now I'm confident that I'll have a job in the future cleaning up slop. Lucky us.
- scandinavian 7mo agoI did a cloc check on it and it does seem to have 800k lines of typescript. So unless they are vendoring dependencies it's actually as insane as it sounds.
- jsheard 7mo agoChrist their repo is an absolute nightmare. There's new issues and PRs being posted practically every minute, and I assume 99% of them are from agents given the target demographic. Just full-auto vibeslop from all barrels 24/7. Even if we count the repos whole lifetime, including when it wasn't so active, the averages are still absurd. 96 days / (4,239+9,170) issues = one issue every 10 minutes 96 days / (5,082+10,221) pull requests = one PR every 9 minutes
- CrazyStat 7mo agoAt least nobody can accuse them of not dogfooding enough.
- mihaelm 7mo ago5000+ open PRs is pretty insane, that's the highest I've seen. How do you even keep track of this? We'll really need trust management systems like vouch (https://github.com/mitchellh/vouch/tree/main https://github.com/mitchellh/vouch/tree/main) for open source projects in the future to help with reducing noise.
- buremba 7mo agoMy take is that agents should only take actions that you can recover from by default. You can gradually give it more permission and build guardrails such as extra LLM auditing, time boxed whitelisted domains etc. That's what I'm experimenting with https://github.com/lobu-ai/lobu https://github.com/lobu-ai/lobu 1. Don't let it send emails from your personal account, only let it draft email and share the link with you. 2. Use incremental snapshots and if agent bricks itself (often does with Openclaw if you give it access to change config) just do /revert to last snapshot. I use VolumeSnapshot for lobu.ai. 3. Don't let your agents see any secret. Swap the placeholder secrets at your gateway and put human in the loop for secrets you care about. 4. Don't let your agents have outbound network directly. It should only talk to your proxy which has strict whitelisted domains. There will be cases the agent needs to talk to different domains and I use time-box limits. (Only allow certain domains for current session 5 minutes and at the end of the session look up all the URLs it accessed.) You can also use tool hooks to audit the calls with LLM to make sure that's not triggered via a prompt injection attack. Last but last least, use proper VMs like Kata Containers and Firecrackers. Not just Docker containers in production.
- alexhans 7mo agoThat's a decent practice from the lens of reducing blast radius. It becomes harder when you start thinking about unattended systems that don't have you in the loop. One problem I'm finding discussion about automation or semi-automation in this space is that there's many different use cases for many different people: a software developer deploying an agent in production vs an economist using Claude Vs a scientist throwing a swarm to deal with common ML exploratory tasks. Many of the recommendations will feel too much or too little complexity for what people need and the fundamentals get lost: intent for design, control, the ability to collaborate if necessary, fast iteration due to an easy feedback loop. AI Evals, sandboxing, observability seem like 3 key pillars to maintain intent in automation but how to help these different audiences be safely productive while fast and speak the same language when they need to product build together is what is mostly occupying my thoughts (and practical tests).
- daveguy 7mo ago
- rdtsc 7mo ago> The container boundary is the hard security layer — the agent can’t escape it regardless of configuration I thought containers were never a proper hard security barrier? It’s barrier so better than not having it, if course.
- rco8786 7mo agoIn the sense that nothing is truly a "proper" hard security barrier outside of maybe airgapping, sure. But containerization is typically a trusted security measure.
- TeeWEE 7mo agoDo you trust your employees? Do you trust a contracter? Do you trust other people? AI is similar to a person you dont know that does work for you. Probably AI is a bit more trustworthy than a random person. But a company, needs to let employees take ownership of their work, and trust them. Allow them to make mistakes. Isnt AI no different?
- TeeWEE 7mo agoMy point is: Trust the work of AI just like the work of a contracter: Check and verify, but dont micromanage.
- adam12 7mo agoCan you sue an ai agent?
- alexhans 7mo agoI think a key ingredient here is accountabilty and liability. If there's a mistake, you can't blame the computer. Who is the human accountable at the end of it all? If there's liability, who pays for it? That's where defining clear boundaries helps you design for your risk profile.
- arnvald 7mo agoIt’s totally different. People have to obey laws and contracts because there are consequences if they don’t, there are fines, arbitrage, courts. What happens if AI agent you run causes a lot of damage? The best you can do is to turn it off
- ramoz 7mo agoYes, it is different. An AI actions and reasons through probabilistic methods - creating a lot more risk than a human with memory, emotions, and rationale thinking. We can’t trust AI to do any sensitive work because they consistently f up. With & without malicious intent, whether it’s a fault of their attention mechanisms, reward hacking, instrumental convergence, etc all very different than what causes most human f ups.
- juggle-anyhow 7mo ago
- ed_mercer 7mo agoHow is Nanoclaw different from running openclaw in a VM?
- xrd 7mo agoHow can I trust this discussion when my browser won't trust their certs?
- badsectoracula 7mo ago> OpenClaw has nearly half a million lines of code, 53 config files, and over 70 dependencies. This breaks the basic premise of open source security. Chromium has 35+ million lines, but you trust Google’s review processes. Most open source projects work the other way: they stay small enough that many eyes can actually review them. Nobody has reviewed OpenClaw’s 400,000 lines. This reminds me of a very common thing posted here (and elsewhere, e.g. Twitter) to promote how good LLMs are and how they're going to take over programming: the number of lines of code they produce. As if every competent programmer suddenly forgot the whole idea of LoC being a terrible metric to measure productivity or -even worse- software quality. Or the idea that software is meant to written to be readable (to water down "Programs are meant to be read by humans and only incidentally for computers to execute" a bit). Or even Bill Gates' infamous "Measuring programming progress by lines of code is like measuring aircraft building progress by weight". Even if you believe that AI will -somehow- take over the whole task completely so that no human will need to read code anymore, there is still the issue that the AIs will need to be able to read that code and AIs are much worse at doing that (especially with their limited context sizes) than generating code, so it still remains a problem to use LoCs as such a measure even if all you care are about the driest "does X do the thing i want?" aspect, ignoring other quality concerns.
- gyomu 7mo agoYeah, it’s pretty wild. Even pg is tweeting stuff like “An experienced programmer told me he's now using AI to generate a thousand lines of code an hour.“ https://x.com/paulg/status/2026739899936944495 https://x.com/paulg/status/2026739899936944495 Like if you had told pg to his face in (pre AI) office hours “I’m producing a thousand lines of code an hour”, I’m pretty sure he’d have laughed and pointed out how pointless that metric was?
- ElProlactin 7mo agoEnshittification comes for us all
- medi8r 7mo agoHe is a Lisper too, making it more ironic. Lisp the power to heavily reduce cruft by heavy customization with macros.
- nemo44x 7mo agoI’ve seen skills, etc haphazardly being launched with no constraints or guardrails. That more or less have admin access and can take actions that are not reversible. It’s the monkey with a gun meme.
- deleted 7mo ago[deleted]
- nkzd 7mo agoAs someone who only coding agents at work, can someone describe their use case for claw type agent? What do you do with it?
- medi8r 7mo agoI want to try one to be a bit of a personal coach. Remind me to do things and check in on goals. The memory / schedule / chat thing is enough and it wont need emails or anything more dangerous.
- theturtletalks 7mo agoHas anyone used: OpenClaw NanoClaw IronClaw PicoClaw ZeroClaw NullClaw Any insights on how they differ and which one is leading the race?
- huqedato 7mo agoThe same crap under the hood, IMO.
- redman25 7mo agoYeah, good software takes time. These are all popping up way to fast.
- tao_oat 7mo agoI haven't used them all but based on my partial research so far: - OpenClaw: the big one, but extremely messy codebase and deployment - NanoClaw: simple, main selling point is that agents spawn their own containers. Personally I don't see why that's preferable to just running the whole thing in a container for single-user purposes - IronClaw: focused on security (tools run in a WASM sandbox, some defenses against prompt injection but idk if they're any good) - PicoClaw: targets low-end machines/Raspberry Pis - ZeroClaw: Claw But In Rust - NanoBot: ~4k lines of Python, easy to understand and modify. This is the one I landed on and have been using Claude to tweak as needed for myself
- theturtletalks 7mo agoWhich would you say has the best cron and heartbeat implementation?
- tao_oat 7mo agoHaven't tried them in enough depth to compare. Nanobot's was not great (cron + a HEARTBEAT.md meant two ways to do things, which would confuse the AI). But because the implementation is so simple, I could improve it in a few minutes in my own fork!
- echoangle 7mo agoLooking at the NanoClaw GitHub README: > If you want to add Telegram support, don't create a PR that adds Telegram alongside WhatsApp. Instead, contribute a skill file (.claude/skills/add-telegram/SKILL.md) that teaches Claude Code how to transform a NanoClaw installation to use Telegram. Why would you want that? You want every user asks the AI to implement the same feature?
- deleted 7mo ago[deleted]
- spacecadet 7mo agoWhy this is posted here and is a revelation for anyone, this many years later is indicative of the times. Good bye.
- Sytten 7mo agoI am a caveman, I don't understand the need for a personal assistant. What are you guys using it for?
- vitto_gioda 7mo agoI only use my own “agent” ("my", because I program it myself, since my needs are different from yours) to retrieve information about the audio I upload to it (from video calls and audio recordings). No others use cases for me
- ramoz 7mo agoIm terrible with email, so its be genuinely helpful for me there. Excited to explore more use as time permits. Very optimistic based on email experience. My next use case is personal notes system.
- andrew_eu 7mo agoI set one up to have a shared chat with my partner about our dog. E.g. schedule reminders, tracking food in a spreadsheet, etc.
- rubslopes 7mo agoTools like OpenClaw have two core capabilities: the ability to rewrite themselves, and the ability to independently figure out how to connect to different services and establish those connections. Yesterday, I was responding to a client ticket about what I knew wasn't a bug. It was something the client had requested themselves. The product is complex, constantly evolving, and has spawned dozens of related Jira tickets over time. So I asked my agent to explore the git history, identify changes to that specific feature, and cross-reference them with comments across the related tickets. Within minutes, I had everything I needed to write a clear response. It even downloaded PDF and DOCX files the client had attached. All of this was possible because my agent is connected to GitHub and Jira, and can clone repos locally since it runs on a VPS. A second example: I was in an online meeting, taking notes as we went. Afterward, I asked the agent to pull the meeting transcript from Fireflies and use it to enrich my notes in Obsidian. I could have also asked it to push my action items straight into Todoist.
- techpulse_x 7mo ago[flagged]
- vitto_gioda 7mo ago"Time to understand 8 minutes" what a non-technical purpose...
- gmerc 7mo agoOh this can be monetized: claw-guard.org/adnet. Another persons trust issues are your business model.
- Yokohiii 7mo agoWhy do people take this article serious? It's just a wall of gibberish trying to make the product look more "secure" then others. It's not. It adds shallow secure looking random junk without tackling the core issues. Which are not solvable obviously.
- justonceokay 7mo agoI have twice encountered a phone tree AI agent saying my problem could not be solved and then ending the call. One was for PayPal fraud and the other was for closing an unused bank account. For right now my trick is to say I have a problem that is more recognizable and mundane to the ai (i .e. lie) and then when I finally get the human just say “oh that was a bunch of hooey here’s what I’m trying to do”. For PayPal that involved asking for help with a business tax that did not exist. For my bank it involved asking to /open/ a new account. Obviously th AI wants to help me open an account, even if my intention is to close one. That will only work for so long but it’s something
- Kiboneu 7mo ago“If you trust the tool then you’re holding it wrong”
- deleted 7mo ago[deleted]
- mathgladiator 7mo agoI was blown away by OpenClaw until I saw the bill. Ultimately, I think of these ecosystems as personal enhancements and AI costs need to come down dramatically for real problem. Worse, however, is the security theater. I would not want to be the operator for any business built with front-line LLM usage based on a yolo'd agent framework. I'm very happy to use these for silo'd components that are well isolated and have reasonable QA processes (and that can even included agents since now we literally have no excuse to not have amazing test coverage). Their niche is going to be back office support, but even that creates risk boundaries that can be insurmountable. A friend of mine had a agent do sudo rm -rf ... wtf. My view is that I want to launch an agent based service, but I'm building a statically typed ecosystem to do so with bounds and extreme limits.
- cyanydeez 7mo agoLook at AI like what search turned into: feed the user anything, even if wrong because not doing so will make your product look weak. Thats what youll find when you try to make these bag-o-words do reasonable things.
- nickdirienzo 7mo agoI tried NanoClaw and love the skill (and container by default) model. But having skills generate new code in my personalized fork feels off to me… I think it’s because eventually the “few thousand auditable lines” idea vanishes with enough skills added? Could skill contributions collapse into only markdown and MCP calls? New features would still be just skills; they’d bring in versioned, open-source MCP servers running inside the same container sandbox. I haven’t tried this (yet) but I think this could keep the flexibility while minimizing skill code stepping on each other.
- atonse 7mo ago> I think it’s because eventually the “few thousand auditable lines” idea vanishes with enough skills added? I just watched a youtube interview with the creator. He actually explains it well. OpenClaw has hundreds of thousands of lines you will never use. For example, if I only use iMessage, I have lots of code (all the other messaging integrations) that will never be used. So the skills model means that you only "generate code" that _you_ specifically ask for. In fact, as I'm explaining this, it feels like "lazy-loading" of code, which is a pretty cool idea. Whereas OpenClaw "eager-loads" all possible code whether you use it or not. And that's appealing enough to me to set it up. I just haven't put it in any time to customize it, etc.
- nickdirienzo 7mo agoI totally get that, and I'm reminded of plugin architectures (e.g. VSCode extensions or browser extensions). Those extensions don't modify the core codepaths for what they integrate with, but still provide new capabilities for only what I want to use. I guess I don't see extensibility, agentic capabilities, and more code safety (and fewer tokens burned on codemods) as mutually exclusive. Not saying you're saying that fwiw.
- desireco42 7mo agoI think you have issue with your security cert.
- bigstrat2003 7mo agoAll this talk about sandboxing and permissions misses the obvious: since you can't trust the agents, don't freaking use them. It is utterly stupid to give an LLM access to run things on your computer, because nothing you do can stop it from hallucinating garbage that harms your system. The whole "agent" craze is the most incredible display of irresponsibility I have ever seen in this industry.
- skeledrew 7mo ago> don't freaking use them You can't tell people that. People see the obvious benefits of using agents, so the many will always take the leap regardless of what detractors say. Continually iterating on the security model and making it all transparent is the way to go.
- Eggpants 7mo agoI’m using this but using gpt-oss-120B instead of a cloud service. It has been eye opening when I realized the LLM is beings used as a compiler. I asked it to add apple iMessage and apple notes support as I I rather have long responses, like write me a program ideas, not fill my iMessage history. The local LLM, which I believe has limited bash training data, does pretty well. For example: I enjoy industrial music and asked it for the tour data of the band KMFDM which returned they will be in Las Vegas in April for a festival(Sick new world). This festival has something like 20 bands most of which I never heard of. I asked nanoclaw to search all of the band list and generate a listing grouped by the type of music they play: Industrial, rap, etc. It did a good job based on bands I do know. I was pleased as I certainly did not want to do 20 band web searches by hand. It’s still at a bar trick level. It gives me hope that an upgraded agent based Siri-like OS component could actually be useful from time to time.
- SignalStackDev 7mo ago[dead]
- raffael_de 7mo ago> OpenClaw has nearly half a million lines of code, 53 config files, and over 70 dependencies. Isn't OpenClaw just ... while(true) { in = read_input(); if(in) { async relay_2_llm(in); } sleep(1.0); } ... and then some?
- simon_void 7mo agonobody trusts AI agents, that's why they are put in a harness. It's just that I additionally belong to the people who don't trust AI agents to always adhere to harnesses either.
- snowhale 7mo ago[dead]
- aerhardt 7mo agoA question I've been asking myself and which I honestly want to put out there - and I apologize in advance, because you will see me repeat it in other threads, out of genuine curiosity: Does your life have so much friction that you need a digital agent to act on your behalf? Some of the use cases I saw on the OpenClaw website, like "checking me into a flight", are non-issues for me. I work in business automation, but paradoxically I don't think too much about annoyances in my private life. Everything feels rather frictionless. In business, I see opportunities to solve friction and that's how I make money, but even then, often there are barriers that are very hard to surmount: (a) problems are complex to solve and require complex solutions such as deterministic or ML systems that LLMs are not even close to being able to create ad-hoc (b) entrenched processes and incumbent organizations create moats that are hard to cross (ex: LinkedIn makes automation very hard) (c) some degree of friction, in some cases, may actually be useful! I imagine there are similar dynamics in the consumer space, but more than anything, I may not be seeing issues with such a critical eye (I like to relax after work, after all) So, do you have problems in your private life that you'd want to take on the risks - and friction - of maintaining these agents?
- ethbr1 7mo agoSimilar CV, similar take. My guess? Anyone involved in automation for >2 years at the enterprise levels knows in their gut all the silent, sudden, annoying ways automation can fail and so has a higher internal bar for "must save this much time to be worth automating." That said, old beliefs should be challenged by new technological capabilities! If LLM based automation is (a) less fragile and (b) quicker to develop, then that bar should be lowered.
- wolvesechoes 7mo agoDo not underestimate the modern marketing and its capability to create needs that didn't exist before. It is not about removing friction, it is about convincing that friction existed in a first place
- RevEng 7mo agoI don't get it either. I even sat down to try it out and see what it's all about, but I can't think of a single thing in my life that I want an agent to automate. Summarizing news? No, I'd rather just read it. Besides, it's hard to say what will or won't be interesting at any given moment. Reply to emails? No, I want to make sure I say what I mean to say and I don't see why I would tell the whole story to an LLM just so it could rephrase it all. Trade stocks? Dear God no! That's a good way to lose my life savings, and if the solution is to just put in a little then what's the point? Every video and post I see talking about all the things they have automated with agents are things I would never want. Most of them describe content farms - look at what's hot on Twitter, generate a video, post to Reddit, etc. Others like preparing a morning summary are neat and all but so what? If my life was so hectic that I needed a personal assistant to take my calls and book my meetings, I'd hire one. Seriously, what is the thing that agents can do that every ordinary person just can't live without?
- andai 7mo agoI move the security boundary one or two layers up: the Unix user (on main machine I run them as a `agent` user, so they can't read or write my files), or even better, just give it a separate machine. (VPSes are now popular for this purpose, as are Mac Minis. My choice is $50 Thinkpad :) That said I am a fan of Nanoclaw, and especially the philosophy of "it should be small enough to understand, modify and extend itself." I think that's a very good idea, for many reasons. The idea of giving different agents access to different subsets of information is interesting. That's the Principle of Least Privilege. That seems like a decent idea. Each individual agent can get prompt injected, but the blast radius is limited to what that specific agent has access to. Still, I find it amusing that people are running this with strict rulesets, in Docker, on a VM, and then they hook it up to their GMail account (and often with random discount LLMs to boot!). It's like, we need to be clear about what the actual threat model is there. It comes down to trust and privacy. You can start by thinking, "if the LLM were perfectly reliable (not susceptible to random error or prompt injection) and perfectly private (running on my own hardware)", what would you be comfortable letting it do. And then you remove these hypothetical perfect qualities one by one to arrive at what we have now: slightly dodgy, moderately prompt-injectable cloud services. Each one changing the picture in a slightly different way. I don't really see a solution to the Security/Privacy <-> Convenience tension, except "wait for them to get smarter" (mostly done) and "accept loss of privacy" (also mostly done, sadly!)
- jswelker 7mo agoAs a fun thought experiment, when people complain about LLMs, I substitute the word "human" or "employee" into the sentence and see if it is equally true. "You can never really trust an LLM!" -> "You can never really trust an employee!" (Every IT department ever.) "LLMs make shit up." -> "Humans make shit up." (Wow very profound insight.)
- RevEng 7mo agoWhile that will always be true, LLMs do it a lot more often and do so with confidence and poise. We have evolved ways to tell if someone is making shit up (which usually works); LLMs subvert this. We are also being sold the idea that these LLMs are some kind of super intelligence which isn't helping matters.
- pipejosh 7mo ago[dead]
- Alex_001 7mo ago[dead]
- blakec 7mo agoThe proxy-based secret injection approach mentioned upthread is solid for network credentials, but it doesn't cover the local attack surface — your SSH keys, GPG keys, AWS credentials sitting in dotfiles. Those are the actual high-value targets for a compromised agent on a dev workstation. I run Claude Code with 84 hooks, and the one I trust most is a macOS Seatbelt (sandbox-exec) wrapper on every Bash tool call. It's about 100 lines of Seatbelt profile that denies read/write to ~/.ssh, ~/.gnupg, ~/.aws, any .env file, and a credentials file I keep. The hook fires on PreToolUse:Bash, so every shell command the agent runs goes through sandbox-exec automatically. The key design choice: Seatbelt operates at the kernel level. The agent can't bypass it by spawning subprocesses, piping through curl, or any other shell trick — the deny rules apply to the entire process tree. Containers give you this too, but the overhead is absurd for a CLI tool you invoke 50 times a day. Seatbelt adds ~2ms of latency. I built it with a dry_run mode (logs violations but doesn't block) and ran it for a week before enforcing. 31 tests verify the sandbox catches attempts to read blocked paths, write to them, and that legitimate operations (git, python, file editing in the project directory) pass through cleanly. The paths to block are in a config file, so it's auditable — you can diff it in code review. And it's composable with other layers: I also run a session drift detector that flags when the agent wanders off-task (cosine similarity against the original prompt embedding, checked every 25 tool calls). None of this solves prompt injection fundamentally, but "the agent physically cannot read my SSH keys regardless of what it's been tricked into doing" is a meaningful property.
- tabs_or_spaces 7mo ago> OpenClaw has nearly half a million lines of code, 53 config files, and over 70 dependencies. This breaks the basic premise of open source security. Chromium has 35+ million lines, but you trust Google’s review processes. Most open source projects work the other way: they stay small enough that many eyes can actually review them. Nobody has reviewed OpenClaw’s 400,000 lines. It was written in weeks with no proper review process. Yeah, but the world rewarded this by making it the fastest growing github project. The author gets on the podcasts, gets the high profile jobs from big tech. I'm more encouraged to do things this way than being security minded about all this. And there's no accountability to this at all. If an agent leaks private data, the user is to blame and not the author. If Google bans your services for using api keys incorrectly, we cast the bad eye towards Google and not the maintainer than enabled and approved it. There's just so much incentive for for "not reading code" and not developing secure code that is just going to get worse over time. This is the hype and the type of engineering that we all allow either by agreeing or by staying silent. I agree with the author, but the world works off a different set of principles than what we're used to. I just see the world blindly trusting agents more.
- dave_meshimize 7mo agoTreating the LLM as an untrusted execution thread at the OS level is probably the only sustainable way to handle agentic autonomy... Most frameworks try to manage permissions with application level logic which is basically just a game of whack a mole with prompt injection.
- adrian-vega 7mo ago[dead]
- Noyra-X 7mo agoThe trust problem is real, and I think the framing of "trust" vs "don't trust" misses the more useful question: trust for what, exactly? In our experience with small business automation, agents work well on bounded, well-defined tasks with clear success criteria and human checkpoints. The failures happen when people deploy agents on ambiguous, open-ended decisions and then walk away. The answer isn't less automation – it's better handoff design between agent and human. What specific failure modes have you seen that you think no amount of prompt engineering or guardrails can fix?
- felix9527 7mo agoThe core problem this article surfaces is forensic: once the agent session ends, the evidence is whatever the vendor chose to log. Terminal scrollback is lossy, session logs are vendor-controlled, and "undo" only works if you catch it in time. Certificate Transparency (RFC 6962) solved a structurally identical problem for TLS certificates after the DigiNotar incident. The insight: commit every action to an append-only Merkle tree where any third party can verify inclusion proofs — without trusting the log operator. Applied to agents: - Inclusion proof: "this specific action was recorded at position N, and the log hasn't been rewritten" - Consistency proof: "between checkpoint A and B, the log only grew — nothing was removed or altered" This gives you verification-based accountability, not trust-based logging. The difference matters: a signed receipt proves the signer said something happened. A Merkle inclusion proof proves it's part of a complete, append-only sequence — deletions are structurally detectable. The AI agent ecosystem is having its DigiNotar moment weekly. We have the cryptographic tools to fix it. The question is whether we'll wait for a catastrophic incident to force adoption, like we did with CT.
- felix9527 7mo agoBeen running Claude Code with hooks recording every tool call for two weeks. ~25k actions across 116 sessions. The thing that surprised me: 24 tool calls per prompt on average. "Just review everything" is not realistic at that ratio. What worked for me was the dashcam approach — don't try to prevent, just have an independent record of what the agent actually did. Not the agent's summary, the actual sequence. Caught a few cases where the summary glossed over failed attempts and retries that mattered.