Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gregfrank
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
gregfrank
7mo ago
"Trendslop" is a great name for something I think is a deeper structural problem than it appears. The issue isn't just that LLMs produce generic outputs, it's that our evaluation methods reward the appearance of the righ
2.
▲
by
gregfrank
7mo ago
This framing points at something important that I think the alignment evaluation literature often misses: the distinction between what a model represents internally and what it does behaviorally. Probing can tell you what's in the repr
3.
▲
by
gregfrank
7mo ago
The "linear" assumption here is worth interrogating. In work I've been doing on alignment evaluation, I find that linear probes can achieve high accuracy on refusal-relevant directions, but that probe accuracy is non-diagnost