4 ms·
> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We beli
by tony_cannistra 6mo ago
> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date. How can these claims all be true at once? Consider the ways in which a careful, seasoned mountaineering guide might put their clients in greater danger than a novice guide, even if that novice guide is more careless: The seasoned guide’s increased skill means that they’ll be hired to lead more difficult climbs, and can also bring their clients to the most dangerous and remote parts of those climbs. These increases in scope and capability can more than cancel out an increase in caution.
https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf#page=53.09 https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...
- tekacs 6mo ago"We want to see risks in the models, so no matter how good the performance and alignment, we’ll see risks, results and reality be damned."
- randomcatuser 6mo agoi mean, to be fair, these are professional researchers. i'm very inclined to trust them on the various ways that models can subtly go wrong, in long-term scenarios for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company? another hot use case: biohacking. if a model is used to do really hardcore synthetic chemistry, one might not realize that it's potentially harmful until too late (ie, the human is splitting up a problem so that no guardrails are triggered)
- cruffle_duffle 6mo ago"for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company?" But who gets to be the judge of that kind of "misalignment"? giant tech companies?
- riwsky 6mo agoMight makes right; brains hold reigns.
- Zee2 6mo agoAlignment “appearing” better as model capabilities increase scares the shit out of me, tbh.
- arcanus 6mo agoConversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!
- falcor84 6mo agoIs that actually well defined given the very low sample size at the top? To the best of my knowledge, none of the individuals believed to have an IQ >200 have committed an actual crime. The closest I found is William James Sidis's arrest for participating in a socialist march.
- RugnirViking 6mo agoIQs more than about 140-150 don't really mean much. They typically come from mathematical extrapolation that tries to account for age (this young child performs very well on the test, just think what they can do when they're an adult). Adult scores usually show this not to be the case
- O5vYtytb 6mo agoIf you're smart enough you just use the laws as written to get what you want, or change them.
- sciencejerk 6mo agoYep
- lelanthran 6mo ago> Conversely: in humans, intelligence is inversely correlated with crime. If you're measuring the intelligence of criminals who have been caught, why would you expect it to be otherwise? IOW, you're recording the intelligence of a specific subset of criminals - those dumb enough to be caught! If you expand your samples to all criminals you'd probably get a different number.
- goekjclo 6mo agoI don't know if they can be any more 'cautious' for Mythos 2...
- CamperBob2 6mo agoTranslation: yay, more paternalism.
- kay_o 6mo agoAnthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world
- wolttam 6mo agoIf there are advancements, they have to be described somehow. What if the capability advancements are real and they warrant a higher level of concern or attention? Are we just going to automatically dismiss them because "bro, you're blowing it up too much" Either way these improvements to capabilities are ratcheting along at about the pace that many people were expecting (and were right to expect). There is no apparent reason they will stop ratcheting along any time soon. The rational approach is probably to start behaving as if models that are as capable as Anthropic says this one is do actually exist (even if you don't believe them on this one). The capabilities will eventually arrive, most likely sooner than we all think, and you don't want to be caught with your pants down.
- kay_o 6mo agoI believe advancements sure. But it is a very boy who cried wolf situation for some of these. There are other companies that behave less in this way, Antrhopic seem very unique in that they love making every single release a world ender
- bloppe 6mo agoAltman called GPT-2 "too dangerous to release". Google tends to be much more measured even though they're the ones who tend to release the actual research breakthroughs
- game_the0ry 6mo agoThere is some unintentional good marketing here -- the model is so good its dangerous. Reminds me of the book 48 Laws of Power -- so good its banned from prisons.
- gpm 6mo agoUnintentional? This sort of marketing has been both Antrhopic's and OpenAI's MO for years...
- FergusArgyll 6mo agoBusiness Negging https://www.lesswrong.com/posts/WACraar4p3o6oF2wD/sam-altman-s-business-negging#:~:text=Quoting%20from%20Matt%20Levine's%20Money%20Stuff%20newsletter:,a%20company%20and%20you%20are%20like%20%E2%80%9C https://www.lesswrong.com/posts/WACraar4p3o6oF2wD/sam-altman...
- mbil 6mo agoAgree. I think they're intentionally sitting on the fence between "These models are the most useful" and "These models are the most dangerous". They want the public and, in turn, regulators to fear the potential of AI so that those regulators will write laws limiting AI development. The laws would be crafted with input from the incumbents to enshrine/protect their moat. I believe they're angling for regulatory capture. On the other hand, the models have to seem amazingly useful so that they're made out to be worth those risks and the fantastic investment they require.
- manmal 6mo agoThey should pick a lane because it’s not very believable if you put these things into defense systems and in the next minute claim that humanity is existentially threatened. Either you’re lying, or ruthless, or stupid.
- bitwize 6mo agoThe new Power Mac® G4 with Velocity Engine®. So powerful, the government classifies it as a supercomputer and a potential weapon.
- m3kw9 6mo agoit was trying to hide what it did from an example fix, so how is that tested for alignment