3 ms·
Juniors do learn though, by contrast (as you said yourself). So these models are, at best, perpetual juniors. Except, they make bizarre mistakes that human jun
by solarwindy 1y ago
Juniors do learn though, by contrast (as you said yourself). So these models are, at best, perpetual juniors.
Except, they make bizarre mistakes that human juniors would not, and which are hard for even experienced developers to spot, because while it becomes possible with experience to preempt the kinds of mistakes that stem from the flawed or incomplete mental models of human juniors, these models do not themselves have a model of the computation underlying the code they produce—which in my experience makes the whole process of working with them endlessly frustrating.
- groby_b 1y ago> Except, they make bizarre mistakes that human juniors would not, You underestimate juniors - Stack Overflow has done marvels to expand the capacity for bizarre mistakes :) > these models do not themselves have a model of the computation underlying the code they produce—which in my experience makes the whole process of working with them endlessly frustrating. I think that's... up for debate. They are able to do work that requires at least a very convincing facsimile of understanding. (My favorite benchmark is a couple of 100-500 line pieces of hand-obfuscated C++ with eldritch template horrors, and "what does this do" elicits a perfectly good explanation - and the code is definitely not in the corpus) What they don't have is an ability to infer constraints you didn't explicitly or implicitly spell out. And only limited capability to ask clarifying question. Therein lies the rub - each interaction requires amounts of detail that you only have at the very beginning of a work relationship with a junior. You _can_ spell it out, and they do much better. It's just extremely tedious. They also are horrible at correcting mistakes through ongoing conversation. If your LLM goes off into error land, just give up, reformulate the entire problem, and try again. In essence, you write a PRD and DD, plus a task breakdown, every single time. If you work on something that justifies this overhead - great time savers. If you work on something that can be done sloppily - great time savers. But a lot of code is in between those two lane markers.
- solarwindy 1y ago> I think that's... up for debate Been trying to inform myself on how these models work, and it’s pretty interesting, I have to say. Came across this paper from Anthropic, Scaling Monosemanticity [0], where they’re extracting features (via a trained sparse autoencoder) from the ‘middle’ layer of Claude 3 for the purpose of interpretability, and quite convincingly find features corresponding to abstract concepts that do seem to encode a model of computation of sorts. Most remarkable to me is their example of a feature that activates for functions implementing addition, which holds up under function composition. I guess there’s more going on under the hood than I’ve been giving credit. Of course, that one example is a tiny window into it, and I recognise that even being able to extract that kind of insight into the model’s workings is a feat. > What they don't have is an ability to infer constraints you didn't explicitly or implicitly spell out. And only limited capability to ask clarifying question. Interesting to think about how the concept of clarification can be formalized and whether it’s possible to work in to the next-token prediction paradigm. I have too many holes in my understanding at this point to go much further with the idea... > They also are horrible at correcting mistakes through ongoing conversation I guess this one is somewhat understandable with how the models work, though it’s unfortunate that the typical chat interface strongly encourages you to attempt to resolve your issue through conversation. > a lot of code is in between those two lane markers Yup. [0]: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html https://transformer-circuits.pub/2024/scaling-monosemanticit...