4 ms·
It's often simply misleading / bad writing. Here's one I just got about some crashes: "If the crashes stop, the factory overclock is marginal; run a small neg
by EarthLaunch 20d ago
It's often simply misleading / bad writing. Here's one I just got about some crashes:
"If the crashes stop, the factory overclock is marginal; run a small negative offset."
This looks like it's saying: "If the crashes stop then we know the factory overclock is marginal." (This makes no sense.)
What it's trying to say is: "If the crashes stop then we can run a small negative offset, because the factory overlock is marginal."
What I would write: "If the crashes stop, we can avoid crashes by underclocking slightly. The speed difference between that and factory clock is marginal."
- r_lee 20d agoI'm guessing it's because the way the first one was written looks real smart and sophisticated, which I'm presuming the models are rewarded for, especially when they're fed all kinds of PhD papers and so on as high quality, high weight data
- derektank 19d agoIs it possible that the first message is more information dense/less likely to be ambiguous than the latter? It’s clearly being selected for for some reason, maybe it’s an artifact of the tokenizer or specific training data, but I don’t know. If the use of jargon was complete cruft, I would expect it to be selected against during reinforcement learning
- cruffle_duffle 19d agoYou’d think that, I thought that… but then I realized I’m just kidding myself thinking its output makes sense. It doesn’t. It doesn’t. Sometimes it might as well just speak tongues. In other words, it ain’t you. It’s the model. It’s just genuinely bad. Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”. But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?
- r_lee 19d agoyeah, I just switched to the GPT models and it's a breath of fresh air. as for the other guy, the claude talk is definitely not less ambiguous, it often is incredibly ambiguous and hard to parse, I have no clue why it produces such output, if not to fingerprint it? it's really weird man. when Opus 5 came out, I was really confused. I saw a bunch of hype about how it's better than fable, but I just felt frustrated with it, although at times it'd do fine, but especially in Claude Code it'd just delve into the whole "load bearing" type of lingo real fast and I'd get a headache. I don't think it's worth using even if it scores 2 points higher in some bs benchmark it's definitely surprising how the magic and smoothness of 4.6 and such is no longer there with the >5 models
- lmz 20d agoI thought it meant "the factory overclock is marginal" in the sense of "borderline unstable"?
- ericd 19d agoThis is my read, too.
- fc417fc802 19d agoSame. I read it as "if the crashes stop [ when we test by reducing the clock ] then we know that the overclock applied by the factory is marginal [ ie it barely passed QC or maybe there wasn't proper QC to begin with ] so running with a small negative offset [ ie what we just tested ] can be expected to fix the problem for good". No idea if my reading is right given all the context I'm missing. Either way it's absolutely shit writing in the same way that golfed code is shit code (except when participating in a code golf competition).
- hgoel 19d agoI wonder if this is a result of them trying to cut token consumption by summarizing their RL training data, or maybe it's from how they anonymize user data for training.
- PSZD 19d agoMarginal - definition 2a: of, relating to, or situated at a margin or border. Succinct and precise; a well crafted sentence. A marginal OC results in unpredictable crashes and can be corrected with a small offset; marginality describes the behavior and explains the solution. Inscrutable clues casually conveyed can now be readily explained, at least, unlike the training data of [silence]. Brevity is the soul of wit, but perhaps also exasperated confusion.
- Pannoniae 19d agoI suspect this happens due to optimising for reasoning... if you insert a few words, it will suddenly start to make more sense. "If the crashes stop, (that means) the factory overclock is marginal; (so) run a small negative offset. (to confirm this hypothesis)" The core thought is basically avoid crashes -> caused by marginal overclock -> apply small -offset to test. Which is exactly the order the sentence is in :P