4 ms·
Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.
by Zee2 6mo ago
Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.
- arcanus 6mo agoConversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!
- falcor84 6mo agoIs that actually well defined given the very low sample size at the top? To the best of my knowledge, none of the individuals believed to have an IQ >200 have committed an actual crime. The closest I found is William James Sidis's arrest for participating in a socialist march.
- RugnirViking 6mo agoIQs more than about 140-150 don't really mean much. They typically come from mathematical extrapolation that tries to account for age (this young child performs very well on the test, just think what they can do when they're an adult). Adult scores usually show this not to be the case
- O5vYtytb 6mo agoIf you're smart enough you just use the laws as written to get what you want, or change them.
- sciencejerk 6mo agoYep
- lelanthran 6mo ago> Conversely: in humans, intelligence is inversely correlated with crime. If you're measuring the intelligence of criminals who have been caught, why would you expect it to be otherwise? IOW, you're recording the intelligence of a specific subset of criminals - those dumb enough to be caught! If you expand your samples to all criminals you'd probably get a different number.
- austinjp 6mo agoIt very much depends on the crime. The truly awful stuff is committed by intelligent people.
- naasking 6mo ago> Conversely: in humans, intelligence is inversely correlated with crime. Inversely correlated with crime that's caught and successfully prosecuted, you mean, because that's what makes up the stats on crime. I think people too often forget that we consider most criminals "dumb" because those who are caught are mostly dumb. Smart "criminals" either don't get caught or have made their unethical actions legal.
- mik09 6mo agoyeah anthropic tries to address this through mechanistic interpretation but not sure they are progressing as fast in that domain as their model development