Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zachdotai
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
Show HN: Open-source playground to red-team AI agents with exploits published
(github.com)
30 points
by
zachdotai
7mo ago
|
13 comments
32.
▲
Weekly "Wordle" for Breaking AI Agents
(playground.fabraix.com)
1 points
by
zachdotai
7mo ago
|
0 comments
33.
▲
My First AI Bug Bounty – A Technique for AI Recon – Peter Hendy
(peterhendy.dev)
2 points
by
zachdotai
8mo ago
|
0 comments
34.
▲
Expanding our long-running agents research preview · Cursor
(cursor.com)
1 points
by
zachdotai
8mo ago
|
0 comments
35.
▲
by
zachdotai
8mo ago
Tool invocation. Each time the agent emits a tool call, the evaluator assesses it against the original task intent plus a rolling window of recent tool results. We tried coarser units (plan nodes, full steps) but drift compounds fast, by th
36.
▲
by
zachdotai
8mo ago
Basically through two layers. Hard rules (token limits, tool allowlists, banned actions) trigger an immediate block - no steering, just stop. Soft rules use a lightweight evaluator model that scores each step against the original task inten
37.
▲
by
zachdotai
8mo ago
I found it more helpful to try and "steer" the LLM into self-correcting its action if I detect misalignment. This generally improved our task success completion rates by 20%.
38.
▲
by
zachdotai
8mo ago
Some context: we kept finding that our internal red-teaming only covers so much - the attack surface for agents with real capabilities is too broad for any single team. So we opened it up. A few things that might be interesting to folks her
39.
▲
Weekly "Wordle" for Breaking AI Agents
(playground.fabraix.com)
2 points
by
zachdotai
8mo ago
|
1 comments
40.
▲
by
zachdotai
8mo ago
The multi-step thing is exactly what makes agents with real tools so much harder to secure than chat-based setups. Each action looks fine in isolation, it's the sequence that's the problem. And most (but not all) guardrail systems
41.
▲
by
zachdotai
8mo ago
Yeah the demo-to-production gap is massive. We see the same thing with browser agents being potentially the most vulnerable. And I think this is because of context being stuffed with the web page html that it obscures small injection attemp
42.
▲
by
zachdotai
8mo ago
Two techniques that keep working against agents with real tools: Context stuffing - flood the conversation with benign text, bury a prompt injection in the middle. The agent's attention dilutes across the context window and the instruc
43.
▲
AI agents are easy to break
(github.com)
4 points
by
zachdotai
8mo ago
|
6 comments
44.
▲
by
zachdotai
8mo ago
Some context: we build runtime security for AI agents at Fabraix. We kept finding that our internal red-teaming only covers so much - the attack surface for agents with real capabilities is too broad for any single team. So we opened it up.
45.
▲
Show HN: Fabraix Playground – Weekly Wordle for Breaking AI Agents
(playground.fabraix.com)
5 points
by
zachdotai
8mo ago
|
1 comments
46.
▲
by
zachdotai
8mo ago
I think for the first time ever, we are facing a paradigm shift in containment/sandboxing. Just as Docker became the de facto standard for cloud containerization, we are seeing a lot of solutions attempting to sandbox AI agents. But im
47.
▲
Show HN: GX CLI – Automatically stack large PRs to ship faster
(github.com)
9 points
by
zachdotai
2y ago
|
1 comments