Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bcampbell88
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
bcampbell88
3mo ago
I think the harness determines how you inspect reasoning. Sometimes it self-exfiltrates into the user response. This would be a touchy topic for Anthropic because this distilled reasoning has been used to train other models at scale. To dir
2.
▲
by
bcampbell88
3mo ago
Naturally, this would be nuanced, but I am curious what others would think? The actual damage comes from data leaving the trust boundary, although we are still testing deterministic solutions for ingress. Obviously, you cannot use more prom
3.
▲
Show HN: Prompt Injection as an Egress Problem
(vaibot.io)
3 points
by
bcampbell88
3mo ago
|
1 comments