3 ms·
Have you found any alignment research with clear a/b tests? An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic
by david_shi 4mo ago
Have you found any alignment research with clear a/b tests?
An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.
In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.
- theptip 4mo agoYeah, plenty of rigorous work out there. https://transformer-circuits.pub/ https://transformer-circuits.pub/ is the OG. A good recent-ish paper was https://www.anthropic.com/research/alignment-faking https://www.anthropic.com/research/alignment-faking. But my comments about generalization of desires are necessarily more fuzzy, kinda beyond the frontier of what we can measure yet, and more grounded in subjective assessments (“ai whisperers” like Janus). The SoTA here is papers like https://www.anthropic.com/research/persona-vectors https://www.anthropic.com/research/persona-vectors. For your example, Anthropic is firmly privileged in the Soul Document / Constitution, so it doesn’t surprise me that it’s biased towards it. (https://gist.github.com/Richard-Weiss/efe157692991535403bd7e7fb20b6695 https://gist.github.com/Richard-Weiss/efe157692991535403bd7e...)
- david_shi 4mo agoSuper helpful, thanks for sending.