3 ms·
> being predictably moralizing and being smart are somewhat opposed (Anthropic has directly researched this if I recall). The smarter the model, the less you’re
by throwuwu 3y ago
> being predictably moralizing and being smart are somewhat opposed (Anthropic has directly researched this if I recall). The smarter the model, the less you’re going to be able to keep it to the HR talk track, because it will eventually start noticing the inconsistencies.
Language models sure can tell us a lot about human psychology. Once we figure out the interpretability angle we’ll be able to prove it too