4 ms·
Regularly trying to use LLMs to debug coding issues has convinced me that we're _nowhere_ close to the kind of AGI some are imagining is right around the corner
by alecbz 1y ago
Regularly trying to use LLMs to debug coding issues has convinced me that we're _nowhere_ close to the kind of AGI some are imagining is right around the corner.
- surgical_fire 1y agoAt least Mother Brain will praise your prompt to generate yet another image in the style of Studio Ghibli as proof that your mind is a tour de force in creativity, and only a borderline genius would ask for such a thing.
- ben_w 1y agoSure, but also the METR study showed the rate of change is t doubles every 7 months where t ~= «duration of human time needed to complete a task, such that SOTA AI can complete same with 50% success»: https://arxiv.org/pdf/2503.14499 https://arxiv.org/pdf/2503.14499 I don't know how long that exponential will continue for, and I have my suspicions that it stops before week-long tasks, but that's the trend-line we're on.
- Pulcinella 1y agoBut will it actually get better or will it just get faster and more power efficient at failing to pair parentheses/braces/brackets/quotes?
- ben_w 1y agoRead the linked METR study please. Or watch the Computerphile video summary/author interview, if you prefer: https://m.youtube.com/watch?v=evSFeqTZdqs https://m.youtube.com/watch?v=evSFeqTZdqs
- alecbz 1y agoOnly skimmed the paper, but I'm not sure how to think about "length of task" as a metric here. The cases I'm thinking about are things that could be solved in a few minutes by someone who knows what the issue is and how to use the tools involved. I spent around two days trying to debug one recent issue. A coworker who was a bit more familiar with the library involved figured it out in an hour or two. But in parallel with that, we also asked the library's author, who immediately identified the issue. I'm not sure how to fit a problem like that into this "duration of human time needed to complete a task" framework.
- conception 1y agoThis is an excellent example of human “context windows” though and it could be the llm could have solved the easy problem with better context engineering. Despite 1M token windows, things still start to get progressively worse after 100k. LLMs would overnight be amazingly better with a reliable 1M window.
- alecbz 1y agoWhat does "better context engineering" mean here? How/why are the existing token windows "unreliable"?
- ben_w 1y agoFair comment. While I think they're trying to cover that by getting experts to solve problems, it is definitely the case that humans learn much faster than current ML approaches, so "expert in one specific library" != "expert in writing software".