4 ms·
> It’s astonishing to me that we have to spell this out, that something as obvious as this needs to be explained to LLM’s at all. Hehe. Yeah, that tendency of
by epidemian 2mo ago
> It’s astonishing to me that we have to spell this out, that something as obvious as this needs to be explained to LLM’s at all.
Hehe. Yeah, that tendency of LLMs to document "the story" of the code instead of its current purpose (or non-obvious implementation details) is a pet peeve of mine too. I've added a slew of guidelines to try to sway Claude to not do this, but it still does it often.
At the same time, it feels like something to be expected to have this "failure mode". The model has its context to work on, and what is on its context if not the conversation you've been having (and its internal monologue) and the files it has read? It makes sense that it references the story on its text generations, because that behavior is usually a good thing for an LLM to do. Otherwise, what would it generate? If it generated things that had nothing to do with the conversation in its context, in many cases those things would be seen as "hallucinations", and they'd tend to be RLHF'ed out. So the models that we end up having are the ones that have been reinforced to be most "contextually relevant" and less "hallucinatory".
I might be completely wrong on that of course. It's just my intuitive reasoning of why this seems to be such a prevalent behavior.
- ninkendo 1mo agoRLHF has the same problem as human reviews of AI code: AI code (and comments) look plausible at first, and if you have 100 other PR's to get to, it looks "good enough" and you approve it. I'm sure the humans doing the "human feedback" part of RLHF at anthropic are just as tired as I am at reading all of it, and start to just approve it when it looks plausible. It's doubly insidious because it trips up the human brain too: When I'm reading a PR saying "fix lock ordering to avoid deadlocks on user deletion", and there's a comment somewhere in the diff saying "// use the fixed lock ordering here", my brain tends to completely forget the fact that the comment doesn't make any sense in its surrounding context. Because it makes perfect sense in the context of being the human reviewing the diff. But it's a slight bit of mental effort to remind yourself "what is this comment going to look like to someone reading the code after this is merged?" I wouldn't be surprised whatsoever if the RLHF supervisors forget to apply that extra bit of mental effort and say "yup, this comment looks great", forgetting to check the surrounding code to see if it makes sense on its own.