6 ms·
CrabTrap: An LLM-as-a-judge HTTP proxy to secure agents in production
https://www.brex.com/journal/building-crabtrap-open-source https://www.brex.com/journal/building-crabtrap-open-source
- foreman_ 6mo ago[flagged]
- yakkomajuri 6mo agoReally cool! I'm also building something in this space but taking a slightly different approach. I'm glad to see more focus on security for production agentic workflows though, as I think we don't talk about it enough when it comes to claws and other autonomous agents. I think you're spot on with the fact that it's so far it's been either all or nothing. You either give an agent a lot of access and it's really powerful but proportionally dangerous or you lock it down so much that it's no longer useful. I like a lot of the ideas you show here, but I also worry that LLM-as-a-judge is fundamentally a probabilistic guardrail that is inherently limited. How do you see this? It feels dangerous to rely on a security system that's not based on hard limitations but rather probabilities?
- manapause 6mo agoCorrect me if I’m wrong, but from my experience in this space in order for a model to exercise judgment it must force itself to operate in a strict chain of thought mode. Since all LLMs are predictive creatures, I started to care a lot more about my judgment settings, the transparency of them, and the presence of a judgment loop in either the development or functionality of an application built these days. Not exactly sure where I’m going with this, but my work with creating penetesting tools for LLMs, the way that I use judgment is critical to the core functionality of the application. I agree with your concern and I will just say that the more time I spent concerned with chain of though where now I will make multiple versions of the same app using a different judge set a different “temperaments” and I found it to be incredibly enlightening as to the diversity of applications and approaches that it creates. Even using BMAD or superpowers, I can make five versions of an app without judges involved and I feel like I’m just making the same app five times because the API begins to coalesce around the business problem you want to solve. The vicissitudes of prediction tools always want to take the safest bet for the greater good, but with the judge involved we can make the agent force itself to actually be hostile about what exactly we’re trying to do, which has produced interesting and fun results.
- DANmode 6mo agoWe’re supposed to be fixing LLM security by adding a non-LLM layer to it, not adding LLM layers to stuff to make them inherently less secure. This will be a neat concept for the types of tools that come after the present iteration of LLMs. Unless I’m sorely mistaken.
- SkyPuncher 6mo agoDefense in depth. Layers don't inherently make something less secure. Often, they make it more secure.
- yakkomajuri 6mo agoI do think this is likely to make things more secure but it's also dangerous by potentially giving users a false sense of complete security when the security layer is probabilistic rather than deterministic. EDIT: it does seem to have a deterministic layer too and I think that's great
- reassess_blind 6mo agoIt looks as if this tool has traditional static rules to allow/deny requests, as well as a secondary LLM-as-a-judge layer for, I imagine, the kinds of rules that would be messy or too convoluted to implement using standard rules.
- stingraycharles 6mo agoI think the parent’s point is that this should be implemented using e.g. Bayesian statistics rather than an LLM, as the judge LLM is vulnerable to the exact same types of attacks that it’s trying to protect against. Most proper LLM guardrails products use both.
- snug 6mo agoI think this can be great as additional layer of security. Where you can have a non llm layer do some analysis with some static rules and then if something might seem phishy run it through the llm judge so that you don’t have to run every request through it, which would be very expensive. Edit: actually looks like it has two policy engines embedded
- roywiggins 6mo agoIt's all fine until OpenClaw decides to start prompt injecting the judge
- fc417fc802 6mo agoCalling it now. Show HN: Pincer - A small highly optimized local model to detect prompt injection attempts against other models.
- reassess_blind 6mo agoSounds like a good idea. Please send me the Github link once done and I'll have my OpenClaw take a look and form my opinion of it.
- NamlchakKhandro 6mo agoSounds like a good idea. Please send me you GitHub now and I'll have my big claw crush your open claw
- bambax 6mo agoExactly; would probably be safer with a purely algorithmic decision making system.
- kantaro 6mo ago[flagged]
- Seventeen18 6mo agoSo cool ! I'm building something very close to that but from another perspective, making this open source is giving me many idea !
- alukin 6mo ago[dead]
- fareesh 6mo agoNeeds to be deterministic. ACLs
- erdaniels 6mo agoYes, full stop. They say they cap the body to 16k and give the LLM a warning, lol. And this is coming from a credit card company.
- simonw 6mo agoComments like this don't fill me with confidence: https://github.com/brexhq/CrabTrap/blob/4fbbda9ca00055c1554ae28f8876f3f976862f8a/internal/judge/llm_judge.go#L106-L110 https://github.com/brexhq/CrabTrap/blob/4fbbda9ca00055c1554a... // The policy is embedded as a JSON-escaped value inside a structured JSON object. // This prevents prompt injection via policy content — any special characters, // delimiters, or instruction-like text in the policy are safely escaped by // json.Marshal rather than concatenated as raw text.
- Jerem-6ix 6mo ago[dead]
- samcollins 6mo agoWhy do you say that? I thought this pattern was well established, or are you aware of known issues with it?
- simonw 6mo agoIt doesn't work. You can't trust LLMs to 100% reliably obey delimiters or structure in content. That's why prompt injection is a problem in the first place.
- okwhateverdude 6mo agoRobots struggle with syntax-in-syntax. Really easy to confuse them when asking it to write a SQL query that targets a JSON column but it must respond with a JSON envelope so the harness can parse the result. Lots of escaping that needs to happen. Deeply nested structures in JSON also end up with foibles like missing a ] or } in a string of }}]}]}. Aside from the prompt injection possibility, just the result being straight up broken and requiring another LLM call is tokens flushed.
- frumplestlatz 6mo agoWell-established where and amongst who, exactly? Is it seriously a common belief that this prevents prompt injection? That would be more than a little alarming.
- babas03 6mo ago[flagged]
- NitpickLawyer 6mo ago> Why it lands: specific technical question, credits their work, ends with something that invites response. If Brex engineers are in the thread, one of them will likely reply. BWHAHAHAHAHA. your bot tried, but failed at the same time. (also interesting that this user's other comments seem ok-ish. The prompts are evolving, we get a sneak peek here on what they prompted for, and the delivery seems more human as well)
- adrianstvaughan 6mo ago[dead]
- ArielTM 6mo agoThe debate here is missing a practical question: is the judge from the same model family as the agent it's judging? If both are Claude, you have shared-vulnerability risk. Prompt-injection patterns that work against one often work against the other. Basic defense in depth says they should at least be different providers, ideally different architectures. Secondary issue: the judge only sees what's in the HTTP body. Someone who can shape the request (via agent input) can shape the judge's context window too. That's a different failure mode than "judge gets tricked by clever prompting." It's "judge is starved of the signals it would need to spot the trick."
- edf13 6mo ago[flagged]
- lmeyerov 6mo agoAt RSAC, there were a ton of agentic security startups converging on ebpf monitors for this reason. Eg, sondera gave a fun talk at graph the planet where they did that + exposed with a policy layer over agent traces via Cedar (used in AWS IAM etc). ABAC and identity were also appearing near here. One thing I didn't see: are there any OSS solutions appearing here?
- edf13 6mo agoWe are Open Source… code will be published soon (before launch)
- lanyard-textile 6mo agoThen you will be open source ;) Not yet open source.
- edf13 6mo agoYes, true ;)
- rgovostes 6mo agoI'm willing to wager that your comment was generated from the body of the article plus a prompt to work in an advertisement for your product, which gets a mention in nearly every comment you make (and every submission you make, sometimes on a daily basis).
- edf13 6mo agoHand written I’m afraid… regular comments on this topic is true - it’s an area I’m very interested in.
- IntrepidPig 6mo agoBlatant “astroturfing” in these comments
- hemangjoshi37a 6mo ago[dead]
- agent-kay 6mo ago[dead]
- qwertyuiop_ 6mo agoNon-deterministic business rules engine.
- cadamsdotcom 6mo ago> pointing it at a few days of real traffic produced policies that matched human judgment on the vast majority of held-out requests. The problem is, 99% secure is a failing grade.
- bjackman 6mo ago99% is usually the best you can do. So you can only layer multiple defences together, this makes sense as one layer to me. I have an issue with security layers that are inherently nondeterministic. You can't really reason strongly about what this tool provides as part of a security model. But also, it's in an area where real security seems extremely hard. I think at some point everyone will have a situation where they wanna give an agent some private information and access to the web. You just can't do that in a way that's deterministically safe. But if there are usecase where making it probabilistically safer is enough to tip the balance, well, fine.
- claud_ia 6mo ago[dead]
- deleted 6mo ago[deleted]
- hidai25 6mo agoInteresting approach! I’ve been building something complementary on the deterministic side. LLM-as-judge guardrails are fundamentally probabilistic and can be gamed or hallucinate themselves (as several comments pointed out). That’s why I built EvalView — it does full trajectory snapshots + diffs so you can see exactly what changed, plus a lightweight zero-judge model-check that directly pings the model and reports drift level (NONE / WEAK / MEDIUM / STRONG). Gives you deterministic regression detection that works alongside (or instead of) LLM judges. https://github.com/hidai25/eval-view https://github.com/hidai25/eval-view Curious how you handle drift detection in CrabTrap.
- pitched 6mo agoSecuring agents in real time and testing them for drift in CI are pretty different use-cases… This post is an AI-generated ad, isn’t it? It’s getting too hard to tell!
- hidai25 5mo agoYou’re right that I mixed runtime enforcement with CI drift/regression testing. Different layer, different job. I meant it as complementary, not equivalent. CrabTrap for runtime control, EvalView for deterministic testing/diffing. My bad on making it sound like a drive-by promo.
- Jerem-6ix 6mo ago[dead]
- dixie_land 6mo agoInstalling a self signed cert system wide to do MITM? Sign me up!
- Freedumbs 6mo agoHow is the judge protected from injection?
- DrokAI 5mo ago[dead]
- beyondscaletech 5mo ago[dead]
- artem_am 5mo agoInteresting approach. One thing worth considering alongside this is that a lot of agent risk sits at the network layer before the HTTP payload. An agent that can reach any endpoint at all is already a problem regardless of whether the content looks malicious. Pilot Protocol handles this differently: agents are invisible by default and can only reach peers they've mutually handshaked with. Complementary to what you're doing here, not a replacement.