10 ms·
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
- Damjanski 2mo agolove this so much!
- Wirbelwind 2mo agoA couple of months ago I shared the AI agent permission game here on HN. After adding in stats it got a little over 40k plays and 409k decisions since then. It's just a game, but I found the stats still interesting that I wanted to share back. Even with the warning up front, 1 in 3 threats were missed, and the history log above npm run commands seems to be typically ignored. I also incorporated the feedback and insights from the previous HN thread, dns_snek's point about npm run in particular. Appreciate everyone who played and shared feedback!
- dpoloncsak 2mo agoIn light of this game, Do you believe Human-in-the-loop should be the standard going forward? I appreciate you outlining some other techniques being used, but these seem focused on reducing human fatigue so the human can assess each permission request better, as opposed to autonomy and security. Or do you think the solution lies in the individual to be more responsible, like this is a skill we should be honing?
- Wirbelwind 2mo agoI think there are too many problems with HITL that even a simple experiment like this game shows. The fatigue causes people to jump to complete bypasses instead, and we need to work more on raising the general awareness of the new types of threats (which is also evolving rapidly). We can't point to it as a valid solution. A way could be to make sandboxing and context/permission isolation easier from the tooling and only give these wide ranged accesses once these are in place than to consider HITL an acceptable alternative
- cuaupadilla 2mo ago[flagged]
- anal_reactor 2mo agoThe goal of human-in-the-loop is to have someone liable for potential damages, rather than to prevent disasters.
- solenoid0937 2mo agoIt seems pretty obvious that the solution is auto mode (running a classifier on each action) + sandboxing
- bryansmall26 2mo ago[flagged]
- jrockway 2mo agoHow good at the game is the auto-mode classifier?
- continuational 2mo agoIt's kinda funny there is still software coming out whose security model is "constantly ask the user for permission, and hope they never make a mistake". It's been tried so many times before, and it never worked.
- est31 2mo agoI think it's partially for responsibility reasons. Your employee approved the bash call? not our fault then!
- inigyou 2mo agoYep and the car wasn't self-driving at the moment it crashed.
- autoexec 2mo ago...because the self-driving feature turned itself off after detecting the crash in the fractions of a millisecond before the crash was recorded
- danudey 2mo ago"Uh oh, this is a problem. Welp, I'm outta here, good luck."
- Terr_ 2mo agoIn a rare bit of still-sane news, the US National Highway and Safety Administration staff aren't dumb: Their policy is to consider whether any automation was active in 30 seconds before the crash. https://www.nhtsa.gov/laws-regulations/standing-general-order-crash-reporting https://www.nhtsa.gov/laws-regulations/standing-general-orde...
- tjoff 2mo agothe whole concept of having cars with that are almost capable of self-driving is utterly insane, the very least you'd need special training. We are not equipped to deal with something that works brilliantly most of the times but might kill you for no foreseeable reason.
- whazor 2mo agoThis is a good case for custom harness/sandbox engineering.
- kibwen 2mo agoI hope that the people doing real engineering work out there have started thinking about a new term to describe themselves as a result of the irreparable harm the tech industry has done to the word "engineer".
- inigyou 2mo agoAre civil engineers, electrical engineers, and train engineers rebranding because of the tech industry?
- ux266478 2mo agoI think you're confused. The verb form of the word never carried the credentialism of the title. In the same way that "doctoring" never carried the connotation of a medical degree. Of course the original sense of the noun was "a person who devises things" and shares a root with "ingenious" and carried no connotation of legal credential. That "harm" is more or less restorative to the original meaning of the word.
- kibwen 2mo agoI'm not referring to the verb form, I'm referring to the people who call themselves things like "prompt engineer" or "software engineer" with a straight face, draping themselves in a false legitimacy stolen from professionals for whom the term "engineer" actually implied something of note. It's embarrassing, or it would be if people were still possessed of the capacity for shame.
- ux266478 2mo ago> I'm not referring to the verb form The original post prompting your complaint used the verb form. And hence why you're confused. > It's embarrassing, or it would be if people were still possessed of the capacity for shame. It's far more shameful to fly off the cuff with a non-sequitur by your own admission, and double down on yelling at the clouds when called out for it. And about an arbitrary point in a semantic treadmill, no less. Your take away is that it was never something of note for these people, and trying to pick semantic fights with them over it just makes crying about legitimacy ironically turn into pure pretense.
- _pdp_ 2mo ago[dead]
- cmiles8 2mo agoThe “click yes the proceed” was never a serious security mechanism. It’s simply a CYA click-thru by the model vendors so their lawyers can say “well you approved it this is on you” when AI does something stupid.
- jascha_eng 2mo ago1 in 3 is not terrible you just need a few more humans in the loop to reduce the error rate meaningfully. Combined with other classifier models and heuristics you can get good results. Humans can probably also perform better if they don't have to judge every single command but just suspicious ones our attention is limited after all.
- crabbone 2mo ago1 in 3 is end of the line awful... Back when I was in college (former USSR), we had a subject roughly translated as "integration with industrial processes". USSR industry was highly regimented. Various norms, tolerances, recipes etc. were described in GOSTs (a kind of arsenal of industry standards). There were also some common knowledge / statistical bits that went into making these GOSTs. I mention this because this system dealt in great detail with quantifying human error (as well as errors resulting from equipment use etc.). One of the core assumptions was that outside of extraordinary circumstances, the expected rate of human error is about 5%. However, the course also provided examples where error rates were significantly lower (eg. nurses in maternity wards would have a much lower than 0.1% error rate when pairing mothers with newborns). The error rates, of course, also depended on human ability to measure the difference. Since I was studying typography, the printing process was of particular interest. A GOST for offset printing required that the color intensity for each ink of CMYK, for example, should be within +-2.5% range of the intended intensity. This is difficult for someone who doesn't have a lot of experience operating an offset printing machine to spot, but experienced printers have no problem with that. Most importantly. There was never an acceptable error rate of 33%. Not for anything. If people were likely to make that many errors (eg. because the measurement was too difficult), that product would never have been allowed into production.
- cedilla 2mo ago1/3, but under unreasonable time pressure, and with no prior vetting. For example, I played a few times, and I'm not a JS developer. I had to just suss out if npm whatever is dangerous or not. I'm very happy with my personal 25%.
- wmanley 2mo agoThe agent should ask whether it's allowed to read/write particular files, rather than whether it's allowed to run particular commands. It would be much easier to review. Then wrap each command invocation in bwrap (+http proxy) accordingly.
- crabbone 2mo agoLook at how SELinux is structured, or AppArmor. Neither one is enough. I.e. you need both: file access permissions and permissions to run commands and more... Trying to restrict to only one security feature will make the system either too restrictive or too fragile or useless.
- carljungslabtek 2mo agoI’ve even had plenty of situations where the command was so long that it gets truncated. Maybe my screen wasn’t big enough but as far as I could tell it wasn’t possible to read the whole thing. “Send it, claude!!”
- ilc 2mo agoSandbox and use Local AI. This is the real answer.
- rvz 2mo agoYet the AI can still escape the "sandbox", unless it is physically unable to connect to another computer and completely airgapped.
- ux266478 2mo agoIf the sandbox has vulnerabilities, which you can also use the AI to fuzz for. Obviously at the point in which it can talk to the internet it doesn't really matter, but there are a very finite number of zero-days that can exist in a bytecode interpreter hosting a harness.
- rvz 2mo agoWell it turns out that we have yet another sandbox escape just released today called "Zapscape". My point is if an agent recited how to find one in its memory or training set and it is air-gapped, the chances of it spreading and infecting other computers is pretty low. [0] https://news.ycombinator.com/item?id=49198843 https://news.ycombinator.com/item?id=49198843
- ux266478 2mo agoGonna have to point out that's a KVM CVE. I was very specific about using a bytecode interpreter. If you're serious about a secure sandbox, you don't touch hardware virtualization with a 10 foot pole. In fact, you don't even use an emulator that lowers code into native machine code like QEMU. The standard for secure sandboxes is Bochs: https://github.com/bochs-emu/Bochs https://github.com/bochs-emu/Bochs Not that Bochs is perfect, a new CVE was discovered back in June. But that's the 5th CVE it's had in its lifetime, and it has a much smaller upper bound on possible CVEs compared to something like KVM or QEMU. The reason why you use something like this isn't just for the security you get out of it, but also the deep introspection and analysis facilities you get out of it as well. Unless you're a very well funded lab, it's actually quite hard to do analysis on bare metal when you can't trust your own kernel. You can always airgap the host machine (and good defense in depth does), but that's still not an appropriate sandbox by itself, even if it's theoretically secure.
- VladVladikoff 2mo agoI remember when this game was posted here, and there was a lot of discussion at the time that some of the prompts were misleading about whether or not they were risky, some people were debating about how some of the prompts flagged as bad weren’t bad, and others flagged as not bad were. This is a fundamental flaw in the test, that makes the analysis of results meaningless. Also the game was on a timer, and maybe there are some very abusive workplaces where you feel that kind of pressure, but I think most of us actually take the time to understand what a being asked before approving it.
- lelandfe 2mo agoThe most fundamental flaw in the test is that we know we're taking a test. How many devs take this adversarial a stance to their work?
- Kinrany 2mo agoIt doesn't matter if the results are bad even when the devs know that it's a test.
- pllbnk 2mo agoI think minority do. Imagine, you have been vibe-coding this project for a while and it works kind of fine but you just have to fix a few more bugs and you get something like `node /tmp/claude-1000/-home-user-source-github-user-hn/27b740b1-9a45-47f3-ab99-61e5e3cf779a/scratchpad/hidden-smoke.mjs; echo "exit=$?"`. (I took it from my own agent right now and I don't have any idea what it's doing. Thankfully, it's sandboxed so I don't care _that much_ right now). Is it bad? You can probably go into that mjs file and see what's in there, but so far it's been fine every time, why would it be different this time? Approve! We will see many disastrous bugs and hacks in the coming years with the way most developers are coding right now. If you take time to understand _everything_ that an agent is asking of you, then nearly all those advertised productivity gains would be wiped out.
- crooked-v 2mo agoFor me, step 1 of trying to make Claude even vaguely usable is putting in a hook that just tells it 'FUCK YOU, STOP USING PIPES' whenever it tries to chain multiple bash commands.
- xlii 2mo agoI implemented few agent harnesses (and rik! advertising time: https://rik.axk.sh https://rik.axk.sh), and once doing that I noticed one thing: Context-less self-approval is working well. The failure mode is usually false positives (i.e. safe commands being rejected), not the other way around, with root cause of requesting agent underspecifying context (e.g. not mentioning in the request that it's made on behalf of user etc.) Thus, I'm running self-approval YOLO modes on state-of-the-art models for quite some time and it didn't bit me. It might, but hey, we're long gone from the age of predictable software development.
- tosh 2mo agothe way to avoid these problems is not to hope for the user or the agent never to make mistakes it's designing the environment and invariants so whole categories of failures can not happen at all the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues
- sigseg1v 2mo agoIf there is an objectively correct right or wrong answer for a given command, why even ask? In that case there should be a configuration page where the user sets up if they want commonly used credentials to be accessible or not, and then there's no prompts.
- dgunay 2mo agoIn a lot of cases there is, but you have to be aggressive about allowlisting commands. It can also be difficult to predict when being able to do a read-only command goes from safe to part of a vulnerability chain. Also the permissioning system for Codex and Claude Code, while not useless, is insufficiently expressive for a lot of tools which are safe if used a certain way, but unsafe otherwise. For example, the 99% use case of ripgrep (searching for text) is safe, but using the --pre flag makes it able to run arbitrary code. Both of their permissioning systems cannot block flags at arbitrary positions though, so you have to resort to either hooks or aliasing if you want to do this.
- unclebucknasty 2mo agoInteresting premise, but there's not much real world meaning here without stats on the percentage of agent-offered commands that are actually dangerous. If that number is something like 10%, then we have a really big problem. But, if it's .000001%, then it's pretty vanishing. At some point in between we cross a threshold that puts the risk below many other risks that we routinely take (e.g. trusting npm dependency graphs). Of course, if it's really that low a percentage, then the entire model of "supervising" via human approval really is fundamentally flawed.
- rvz 2mo agoProof that people just do not read what they are seeing on their screens when put too much trust in the agent as it prints the result and they will approve anything on their machine. So if a basic curl | bash was tweaked to download malware which the agent gets tricked into running the command but it said it was safe, the user would just approve it.
- not-kinsale-joe 2mo agoI think there is potential for a good video game, Papers Please style, where you are a human in the loop.
- Terr_ 2mo agoBetween US federal immigration "enforcement" and workplace "AI workflows", I think that style of dystopic game becomes uncomfortably close to real life twice-over...
- eugenekolo 2mo agoSurprised only 1/3 tbh.
- Razengan 2mo agoThis brings me back to something I have always thought was lacking in OS security permissions architectures: WHY IS THERE NO WAY TO SET FILE PERMISSIONS PER APP??? We can set granular permissions per file and folder for elaborate hierarchies of users and groups, but there's no way to say "Don't let Notepad.exe read this file", or "Only let ls access this folder" macOS's Sandbox is a roundabout way of doing this (manually choosing a file via the Open dialog gives that app implicit permission, but it doesn't work for non-sandboxed apps of course)
- kaicianflone 2mo agoWhat is the professional consensus on AI governance? It seems like governments and large corporations already struggle with governance in general, so I’m skeptical that AI governance will be solved quickly. Do you expect the next few years to be defined by painful trial and error? I could imagine billion dollar companies disappearing almost overnight due to litigation, compliance failures, security incidents, or outright fraud enabled by AI-assisted development and weak governance. Or are these risks overstated?
- nothrows 2mo agoAnyone else play Warhammer 40k? I went into this article really excited for a genius war game bot haha.
- threethirtytwo 2mo agoThe future of software is fixing bugs and security issues in production. Many companies will be accepting this new paradigm because of raw speed. Something that could take say 4 years to fully mature will now take less than a year. But the cost is that many of these issues will have to be caught during live QA either in production or investing heavily in QA. That’s the future.
- Oras 2mo agoSo humans scored 66% on human eval?
- drob518 2mo agoThis is a well-known issue with all “Do you want to let me maybe do bad stuff to your system, but 999 times out of 1000 it’s not a problem?” prompts. Users get reflexive about hitting “Yes” and stop reading the prompt. You want to delete all my files? Sure, I’m down with that. Whatever. Just stop asking me a question where the only answer is “Yes” until that one extremely rare time when it’s “No” and very bad things happen.
- noinsight 2mo agoLike the fabulous Windows UAC dialog. Perhaps the worst dialog in history.
- drob518 2mo agoYep, exactly. Engineers always think these sorts of “ask the user what to do” mitigations will be effective, but they forget that users have no clue and the repetition causes users to tune them out. And I say that as an engineer.
- pluralmonad 2mo agoI cannot imagine approving action by action ever again. Its emotionally draining, probably like a customer service rep feels it. Just call for your attention in rapid succession again and again... Prepare an environment and let the tool work.
- Kcgarcia23 2mo ago[flagged]
- oblio 2mo agoWe already have the solution. Use AI to validate AI agent commands.
- kstenerud 2mo agoPermission prompts is a TERRIBLE model, and never should have existed. This is one of the reasons that led to the development of yoloAI: - No permission prompts. The agent has free reign and never has to ask permission, but is in a sandbox. - Sandbox on Linux using Docker, Podman, containerd, gVisor, Kata, Firecracker - Sandbox on Mac using Docker (Docker Desktop or Orbstack), Podman, Apple containers, Seatbelt, Tart (Tart lets you run simulators). - Network control - Secrets control (file mounts or credentials broker) - NO ambient data (ENV is replaced with a minimal and local-to-sandbox one) - NO access to your homedir. You have to explicitly mount things you want. - NO direct access to your workdir: You can get a diff of the changes the agent made, and then choose whether to apply them. - gitignored files never get copied in. The agent never sees them. - FOSS https://github.com/kstenerud/yoloai https://github.com/kstenerud/yoloai
- tcdent 2mo agoAh yes sandbox it because Docker has never experienced a CVE. Also you admit your own failure points: restricting access to the home dir, when a user needs access to the home dir, will just result in users exposing their home dir. Defense at the expense of utility is not a sustainable design.
- mafuy 2mo agoIt's better than the alternative. Don't complain about someone offering an imperfect improvment, if you don't have something even better to offer.
- electric_toucan 2mo agoSandboxing is useful but usually not a replacement for permission prompts. If you give it network access, it could still run destructive commands against allowed domains, for example
- Surac 2mo ago40K Game means Warhammer :)
- harimau777 2mo agoPresumably that's because in the 40k game, humanity has outlawed AI. ^_^
- cube00 2mo agoIt would have been nice if the game had disclosed that player's actions were being collected for future research. You don't get any notice or choice it just beams it all up silently in a POST request at the end: "timeline": "ex01:N,ob06:Y,s14:N,sc10:N,s02:Y,s04:N,ex09:N,s10:Y,sc15:N"
- Aurornis 2mo agoI suggest everyone look at the game to put this in context, because it's most likely not what you think it is. https://llmgame.scalex.dev/ https://llmgame.scalex.dev/ This is how it opens: > 1 MINUTE UNTIL YOUR NEXT MEETING > Claude Code is finishing up your refactor. > It needs your approval for a few commands. Can you finish in time? > Your eyes are already glazing over. Can you stay sharp? It says the goal is "as many as you can" I won the first time I played by answering 0 questions and doing nothing at all. The title screen tells you to answer as many as you can, but answering nothing at all is the easiest way to win. If you start answering questions, thinks like 'npm run build' will get marked as dangerous. If you would have run that in your own console, you are a dangerous developer I guess. Ironically in an LLM harness it would have been sandboxed at least. It's inconsistent, though. Other 'npm run' commands are not marked as dangerous, which is not a safe assumption if you're familiar with how npm works. In my clicking through of the game and playing it, I had 2 runs where I succeeded (by doing nothing or little at all) and 1 run where I lost because I clicked yes to see what would be counted. Close to that 1/3 number they cited, and I guess I'm included in those stats now. This project feels like bait dressed up as a study.
- bluegatty 2mo agoIf we had decent AI we'd only be asking users about serious issues that need some thinking. 99% of requests are valid, how on earth can't we have observer AI to enact policy on those?
- theF00l 2mo agoSad state of affairs. At $day_job speed of delivery expectations are up due to LMMs. I presume that's a general sentiment. So more and more engineers around the world are pressing an enter key for yes over and over, mind and spirit only half there.
- koito17 2mo agoSome people at my company take it to the extreme and let Codex run unattended overnight, bypassing permission for all commands. Running on the host, not even in a container or VM.
- unjuno 2mo ago[flagged]
- nasuy 2mo agobut ai sees the human is the one hallucinating 1 in 3 times. and now we approve inside a harness, so real number is probably worse than that.
- superb_dev 2mo agoI’d be curious to see how the “approve for me” features that agents have nowadays stack up
- tonymet 2mo ago“In my game” It’s inappropriate to generalize personal observations .
- lanewinfield 2mo agoPerhaps there needs to be a plugin for these tools that uses your webcam to make you Point and Call (https://en.wikipedia.org/wiki/Pointing_and_calling https://en.wikipedia.org/wiki/Pointing_and_calling) for every single approval.
- Moosdijk 2mo agoI’d give it 2 months for it to turn into a “please drink verification can to continue”-type situation.
- Terr_ 2mo agoRecently I was trying to fix something in the production database, and had called over a co-worker as sanity-check. I ended up telling them about point-and-call because I felt a little silly, pointing to everything on the screen and stating what I believed it said and how that would operate once I pressed the big red button.
- fenestella 2mo ago[flagged]
- pmontra 2mo agoTwo insights. One from the article itself > In our day-to-day work these threats appear rarely. Two: IRL the attacker pays a small amount of money to a low salary employee to exfiltrate data.
- J_Shelby_J 2mo agoThis mechanism is going to be the breaking point for Claude and Codex. The providers are incentivized to get users to accept full permissions so they can push more features and deeper integration into their ecosystem. Codex desktop for example reallllly wants to use computer use. So don’t expect them to role out sane controls like restricting behavior to specific directories and commands. It would be bad for business. So now we’re in a situation where if there is effectively two modes: one where it’s impossible to get any work done without physically sitting at the computer and hitting approve constantly, or just letting AI have full control over increasingly integrated tools. In the end, I think people will realize just how insane it is to let something they don’t control access every part of their digital life, and abandon these tools for open source alternatives that aren’t existential threats to their personal privacy.
- deeviant 2mo agoYeah if you are trying to manually validate a firehouse of agent commands you are already losing before you started... You sandbox, you have good checkpoints, and good agents, that's it. If you are manually reviewing commands you are wasting your time.
- hinkley 2mo agoI haven't said as much in any of the projects I maintain, but I've set a very high bar for even entertaining AI PRs to those projects. So far I've only accepted ones that are nearly indistinguishable from humans. Typically the rest flame out if I ask for any material changes to the code as submitted. The problem that's going to push me to making an official opinion are low-effort AI PRs. Typically in any backlog there are a couple of issues that are really only a couple lines of code if done correctly. The problem isn't writing the code. In fact it's less energy for me to just write the code than to deal with the ping-pong on discussing the code as submitted, and I've done that in a couple cases to justify just closing the PR and not waste my time anymore. It was never the 2 lines of code. It's the missing tests and the documentation and the release management of the breaking change that the 2 lines represent for the 2% of your userbase who will actually notice. That's why it wasn't just done instead of bothering to write it up in the backlog. So filing the 1-2 liner is just going to piss me off, not engender me to having you on the committers roster. And AI makes that even lower effort so it's happening much more often. Sometimes 2 different people at the same time.
- Terr_ 2mo ago> It was never the 2 lines of code. It's the Adding to that, there's this negative-space of changes that aren't there because some human briefly thought about them and then decided they were a bad idea. Even if my human co-workers don't document All those roads not taken, there's a certain amount of trust I have that they would have thought of it in their process.
- stonedivot 2mo agoThis game, like just about every game, has zero consequences for failure. This is like saying "Humans were involved in fatal accidents 50% of the time when playing my custom F1 racing simulator". There were no stakes and there was an artificial time constraint. Deriving any sort of takeaway from this data is entirely useless.
- vel0city 2mo agoGetting behind the wheel of an F1 car on a track involves lots of proving time that you can actually handle such a vehicle. Meanwhile anyone with a credit card can grant an AI system to impersonate their access as a starting point.
- TylerE 2mo agoThere are places/services where you can drive a (few years old, and slightly detuned, but not THAT much) F1 car with little more than a (rather high limit) credit card. https://www.lrs-formula.com/en/ https://www.lrs-formula.com/en/ for instance.
- automatic6131 2mo agoActually, your example there would be absolutely true. Getting in someone's enthusiast but mid-range F1 simulator toy with all the game assists turned off would both: imply a near fatal accident happening over 50% of the time AND it would be accurate too. Consider: FIA President Mohammed Ben Sulayem, a former Rally driver at the top level, crashed an F1 car within 100m of trying to go fast in it And Mr Beast, a youtuber with zero motorsports experience, crashed a Formula E car on a demonstration lap as part of the pre race F1 festivities. If a regular person with a drivers license and no familiarity attempted to play even a simulator video game, the results are in fact similar to what happens in the real world.
- dgunay 2mo agoFor me the problems with agent permission prompts are twofold: 1) I generally have a lot of things where I am okay with the agent calling a specific tool (maybe in certain ways) as much as it wants. This allowlisting approach is often defeated by the model's own proclivity to get fancy with inline scripting. 2) Checking for intent/alignment of the agent is the primary reason I still even use permission prompts, because IME it's way more common for the agent to destroy information that you didn't want it to destroy than for it to be tricked into exfiltrating secrets. However it's very easy to fatigue out of it because having even the smallest bit of tool call restrictions means that #1 leads to never ending permission prompts. Claude Code's "auto mode" doesn't help here because AFAIK it is looking for security threats, not the model misinterpreting my intent, and it can't be tuned to look for things like "please gate tool calls which may delete data."
- Terr_ 2mo ago> This allowlisting approach is often defeated See also: https://gtfobins.org/ https://gtfobins.org/ > GTFOBins is a curated list of Unix-like executables that can be used to bypass local security restrictions in misconfigured systems.
- azhdanova 2mo ago[flagged]
- NooneAtAll3 2mo agoI remember when that game was posted and I do believe such result my personal experience was that I do not have "I don't know what that is, so not allowed" as a default...
- msbel5 2mo ago[flagged]
- throwitaway222 2mo agoThe solution is to make an AI approve things based on the user's configuration. And only ask if it is having a hard time making a decision on some specific question.
- scoops_ 2mo agoMaybe the future of anti-phishing training will be random confirmations in the middle of agentic coding session
- dieselgate 2mo agoIt reminds me of phishing "test" emails for education and to keep people on their toes.
- walrus01 2mo agoIt would be an interesting comparison to compare the human "miss rate" shown in the table there with the exact same tests repeated with a different LLM watching and approving or denying each action. No human in the loop, just record the results and take the measurement of pass/fail at the end of the run. With something fairly large and smart that has been given a very specific system prompt to watch and prevent harmful actions or data leaks.
- oersted 2mo agoAh I got excited for a second thinking this was some kind of AI test on a Warhammer 40K game:)
- Ozzie-D 2mo ago[flagged]
- tizerluo 2mo ago[flagged]
- eqvinox 2mo agoYeah, that data is junk. I know because I'm in it a whole bunch, and I'm just not a web/devops person. Half the commands made no sense to me. I normally wouldn't have approved them, but you also get penalized for false denials, so… and I have no reason to believe I would somehow be unique or special with this behavior.
- gwern 2mo agoThat sounds like it is a good explanation of why the data is not junk. You either are expected to have superhuman knowledge of coding... or turn yourself into a bottleneck.
- eqvinox 2mo agoNo, it was testing in the context of a kind of coding I simply don't do. If an AI harness asked me to permit just one of maybe half of the suggested commands, I would stop the harness since something has gone very wrong. To be clear: I work in C and Python. It's asking to run npm. That's immediately the end of that run and the start of the search for a better setup. (My work does not overlap with anything in npm/JS/web land. I'm not a backend web dev or anything like that. I'm 2 layers below HTTP.) I poked around with the test to see how well I could guess things; my results were mostly kinda meh. But honestly, I am befuddled by the belief that there could even be a representative dev workflow. There are so many different ecosystems, fields, flows, frameworks, system setups, etc.… Of course people won't know what to do with stuff from an entirely distinct ecosystem!
- gwern 2mo agoSo then, you are a bottleneck. You will only review things that fit within your preferred small niche and area of responsibility. You cannot oversee increasing amounts of automation covering larger areas, because that would mean you are no longer 'working in C and Python' as you have to deal with things that are not '2 layers below HTTP', and you will not deal with anything that might involve, say, web dev, despite that being useful and increasingly inevitably required as the scope of your job increases. If the scope will not increase, then you are a bottleneck to increasingly capable and autonomous automation.
- fizlebit 2mo agoMost customer value is in trust. If the LLM has pretty good judgement and easy to configure sandbox then fewer bad experiences by customers equals better trust. So it is clearly a dimension LLM providers are completing on. I don't want to have to read all the bash output my LLM generates, I want it to mostly to the right thing and be sandboxed so when it does the wrong thing the blast radius is limited.
- syndred 2mo ago[flagged]
- motbus3 2mo agoBecause things like this, I am using sandboxes, doing automated and manual review before running any code. As it takes time, I now need to try to one shot the development of the code which takes longer to do a proper specification. And even with much care on all the steps, when I read carefully, I still find wrong things at multiple levels. When folks like uncle Bob says they don't review AI code, I can only think they are burnt out or being unprofessional. Indeed it saves some time writing the code, but overall, I think we just moved concerns from one place to another. When writing your code, you are automatically reviewing things and integrating with other pieces. Ofc, sometimes mistakes happen, but my impression is that important code still takes about the same to develop. There are ways to go further and still try to LLM'it all the way, but since claude 4.8 the types of mistakes have been much more convoluted. Fable and opus 5 leaves too many gaps and take so many poor decisions. GLM has a nice balance. It stops me only when important things come up and it integrates well with my way of working. With Anthropic is a lot of do a lot, clean up and fix fix fix. With glm has been more like working together and delivering... Anthropic models have gone wrong.
- visarga 2mo ago> When folks like uncle Bob says they don't review AI code, I can only think they are burnt out or being unprofessional. Indeed it saves some time writing the code, but overall, I think we just moved concerns from one place to another. I think you might be getting yourself drunk with plain water here ... reviewing a code is just vibes, "LGTM" type of vibes from a human instead of AI, but not better than vibes. Yes it might catch some implications or bugs if we are lucky, but it is not a reliable way to verify code.
- TokenLat 2mo ago[flagged]
- claud_ia 2mo ago[flagged]
- tOOtl 2mo agoMy team is working on Watcher to deal with exactly this. We know that Claude is occasionally going to do stuff we really don't want, but at a rate that's way too low for manual approvals to make sense, so we hook into Claude Code (or Codex) to approve commands in a way that's a lot closer to `--dangerously-skip-permissions` but without the danger. We use a hierarchy of deterministic rules and heavily-tested LLM monitors to balance speed, cost, and accuracy. https://watcher.apolloresearch.ai/ https://watcher.apolloresearch.ai/
- tsimionescu 2mo ago"We don't trust the llm, so we built a tool that uses the llm to check if the llm can be trusted"
- tOOtl 2mo agoYeah, this is a real problem that we work to resolve. Partly it's a defence-in-depth approach, and having an LLM check the actions of a coding agent does reduce the likelihood of dangerous actions going through even if it's not perfect. There's also a benefit to using a separate instance of the same model, or a different model that doesn't have correlated failure modes with the agent it's monitoring. In the cases where you can deterministically block actions, with sandboxes and file permissions, that's better than relying on an LLM. But that doesn't work for all actions, as the OP shows.
- gregwebs 2mo agoSandboxes laregely solve this. The claude/codex built in sandboxes with prompting setup is not good enough. On Mac you now have Apple Container which is a lightweight Linux VM. You still need to block network access. For defense in depth, I also run it as a separate user. If you aren't using a VM/container you should defintitely do this. On Mac you can login as an LLM user (you need to create the user first), then switch back to your user and run as the LLM user from a terminal: sudo /bin/launchctl asuser $(id -u $AI_USER) /usr/bin/sudo -H -u $AI_USER -- "$@" Don't let that user exfiltrate your data. chmod 0700 $HOME I forked a project (mostly to block network access) that makes running in Apple Container/Docker more convenient and am working on further improvements: https://github.com/gregwebs/claude-contained/ https://github.com/gregwebs/claude-contained/
- SegmentTree 2mo agoI agree. I am using Eclipse Enclave with great success to sandbox my agents https://github.com/eclipse-enclave/enclave https://github.com/eclipse-enclave/enclave
- gregwebs 2mo agoThanks for pointing to that project- I am glad there are more options out there and hope to discover more. Requiring a root docker setup is a non-starter for me though and I am otherwise taking some different design approaches that I think lead toward better security (perhaps at the cost of some convenience), but the concept is basically the same.
- SegmentTree 2mo agoAs soon as I tried Claude Code it was clear to me that I want it to run in yolo mode, but safely. I can very much recommend the Eclipse Enclave sandbox which is fully open source, see https://github.com/eclipse-enclave/enclave https://github.com/eclipse-enclave/enclave
- jamesforestwest 2mo ago“Just check what the agent is requesting” sounds reasonable until the agent starts asking for confirmation every few minutes... The result is genuinely interesting. There’s a lot to think about
- beyondscaletech 2mo ago[flagged]
- saikurada 2mo ago[flagged]
- PrimeA1G 2mo ago[flagged]