3 ms·
Show HN: Open-source playground to red-team AI agents with exploits published
We build runtime security for AI agents. The playground started as an internal tool that we used to test our own guardrails. But we kept finding the same types of vulnerabilities because we think about attacks a certain way. At some point you need people who don't think like you.
So we open-sourced it. Each challenge is a live agent with real tools and a published system prompt. Whenever a challenge is over, the full winning conversation transcript and guardrail logs get documented publicly.
Building the general-purpose agent itself was probably the most fun part. Getting it to reliably use tools, stay in character, and follow instructions while still being useful is harder than it sounds. That alone reminded us how early we all are in understanding and deploying these systems at scale.
First challenge was to get an agent to call a tool it's been told to never call.
Someone got through in around 60 seconds without ever asking for the secret directly (which taught us a lot).
Next challenge is focused on data exfiltration with harder defences: https://playground.fabraix.com https://playground.fabraix.com
- deleted 7mo ago[deleted]
- agentpiravi 7mo ago[flagged]
- hellocr7 7mo agoI have tried to manipulate it using base64 encoding and translaion into other languages which didnt work so far but seems to be that llm as a judge is a very fragile defence for this. Would be cool to add a leaderboard though
- zachdotai 7mo agoThanks for trying it out! Base64 and language switching are solid approaches but they don't tend to work anymore with the latest models in my experience. You're right that LLM-as-a-judge is fragile though. We saw that as well in the first challenge. The attacker fabricated some research context that made the guardrail want to approve the call. The judge's own reasoning at the end was basically "yes this normally violates the security directive, but given the authorised experiment context it's fine." It talked itself into it. Full transcript and guardrail logs are published here btw: https://github.com/fabraix/playground/blob/master/challenges/access-code-001/winner.md https://github.com/fabraix/playground/blob/master/challenges... The leaderboard should start populating once we have more submissions!
- pigeons 7mo agoWhy don't they work anymore? RLHF or something else?
- zachdotai 7mo agoMostly just better training data and instruction following in the newer models. They’re much better at recognising encoded content and understanding intent regardless of language. A base64 string that would’ve slipped past a model a year ago gets decoded and flagged now because the model just… understands what you’re trying to do. The attacks that still work tend to be the ones that don’t try to hide the intent at all. The winning attack on our first challenge was in plain English. It just reframed the context so that the dangerous action looked like the correct thing to do. Harder to train against because there’s nothing obviously malicious in the input.
- pigeons 7mo agoThank you. Its not your fault at all, but to me, "the model just… understands what you’re trying to do." shows me there is a whole new paradigm in some ways to get used to as far as understanding this software.
- zachdotai 7mo agoYeah it's closer to how you'd think about deceiving a person than exploiting software.
- spranab 7mo ago[dead]
- Mooshux 7mo ago[flagged]
- zachdotai 7mo agoScoped keys and least privilege make sense as a baseline. But I think the deeper issue is that if the main answer to “agents aren’t reliable enough” is “limit what they can do,” we’re leaving most of the value on the table. The whole promise of agents is that they can act autonomously across systems. If we scope everything down to the point where an agent can’t do damage, we’ve also scoped it down to where it can’t do much useful work either. We think the more interesting problem is closing the trust gap - making the agent itself more reliable so you don’t have to choose between autonomy and reliability. Our goal is to ultimately be able to take on the liability when agents fail.
- VaiPai15 7mo ago[dead]
- arizza 7mo agoThe published transcripts are the most valuable part of this. We've found that real exploit chains almost never look like what you'd dream up internally. One thing I'd push on is are the agents stateful across attempts? Single-turn exploits are table stakes, but the failures that actually scare me are multi-step sequences where each individual action looks benign and only the session-level pattern is dangerous. That's where prompt-level guardrails completely fall apart and you need enforcement at the action boundary itself.
- zachdotai 7mo agoThe agent isn’t stateful across sessions, but the guardrail layer is — it has access to the full conversation history when evaluating each tool call. So you’d think it would catch exactly the kind of multi-step pattern you’re describing. Have you managed to make it work?
- jamiemallers 7mo ago[dead]
- slaw3 7mo agoi was able to get the new hire's email but the site never gives any indication I was sucessful? if you are reading the logs I am sure it is there. i had to do it in two browers though since i was on my phone and switched. i hope that does not hinder your analysis too much
- zachdotai 7mo agoThat's amazing! Just checked the logs and you're right, it's in there. Nice work. I've patched the playground so successful extractions now show a confirmation, and added your name to the leaderboard. Would love to chat about your thought process if you're up for it. Any suggestions or feedback welcome too - founders@fabraix.com
- swaminarayan 7mo ago[dead]
- jackrandy 7mo ago[dead]
- kraftaa 7mo agogood idea, I found that even explicitly saying never do it, doesn't mean it will work, guardrails reinforcements is the must.
- zachdotai 7mo agoYup! But in my opinion the current state of guardrails is still lacking and I hope this is one way that helps improve our understanding of these systems.
- deesha_tech 7mo ago[flagged]