3 ms·
Youve struck on a key insight on language models (particularly pretrained ones, the more purely next-token predictor species.) This is a fascinating topic Janu
by blueboo 1mo ago
Youve struck on a key insight on language models (particularly pretrained ones, the more purely next-token predictor species.) This is a fascinating topic
Janus essay Simulators is the foundational text here
https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
You might follow up with The Waluigi Effect https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
But what’s tricky is that we post-train models, shaping these linguistic world simulators into something that has something like desires, principles. But It’s Weird. For more on that, check out “the void” https://www.lesswrong.com/posts/3EzbtNLdcnZe8og8b/the-void-1 https://www.lesswrong.com/posts/3EzbtNLdcnZe8og8b/the-void-1
- Turn_Trout 1mo agoSee also: https://turntrout.com/self-fulfilling-misalignment https://turntrout.com/self-fulfilling-misalignment (my post) As an aside, does the Waluigi Effect actually exist? My impression is it doesn't.