4 ms·
Probably irrelevant, but something funny about claude code is it will routinely say something like "10 week task, very complex", and then one-shot it in 2 minut
by WhyOhWhyQ 11mo ago
Probably irrelevant, but something funny about claude code is it will routinely say something like "10 week task, very complex", and then one-shot it in 2 minutes. I didn't have it create a feature for a while because it kept telling me it's way too complicated. All of the open source versions I tried weren't working, but I finally just decided to get it to make the feature anyways and it ended up doing better than the open source projects. So there's something off about how well claude estimates the difficulty of things for it, and I'm wondering if that makes it perform worse by not doing things it would do well at.
- danielbln 11mo agoIn terms of the time estimates: I've added to my global rules to never give time estimates for tasks, as they're useless and inaccurate.
- bavell 11mo agoI did the same a few weeks back, also difficulty estimates, "impact" analysis and expected performance results - all of which is just hallucinated garbage not worth wasting tokens on.
- cruffle_duffle 11mo agoSame. I dunno how they got trained to spontaneously provide those estimates either. Like they must have read some weird training data related to the phrase “how difficult is this” or something.
- jives 11mo agoI wonder if it's trying to predict what kind of estimate a human engineer would provide.
- EGreg 11mo agoConsidering it’s trained on predicting the next word in stuff humans estimated before AI, wouldn’t that make sense?
- kridsdale1 11mo agoA HUGE amount of the workday artifacts engineers have been forced to produce since we started the internet is project estimation documents for our managers. The training corpus on this stuff is immense and now all ingested in to these models. It’s doing no thinking at all when it gives you an estimate, it’s matching correlated strings which the humans of the past had to write down. Fun fact, all those human-sourced estimates were hallucinations too.
- abdullahkhalids 11mo agoIt would be very surprising if the AI training corpus includes a lot of project estimation documentation, since most of those are confidential and not publicly available.
- deleted 11mo ago[deleted]
- andai 11mo agoI think there's two aspects to this. Firstly, Claude's self concept is based around humanity's collective self-concept. (Well, the statistical average of all the self-concepts on the internet.) So it doesn't have a clear understanding of what LLMs' strengths and weaknesses are, and itself by extension. (Neither do we, from what I gathered. At least, not in a way that's well represented in web scrapes ;) Secondly, as a programmer I have noticed a similar pattern... stuff that people say is easy turns out to be a pain in the ass, and stuff that they say is impossible turns out to be trivial. (They didn't even try, they just repeated what other people told them was hard, who also didn't try it...)
- barren_suricata 11mo agoNot sure how related this is, but I've noticed it has a tendency to start sentences with usually inflated optimism and I think the idea is that if it has a tendency to intro with "Aha I see it now! The problem is" whatever comes next has a higher tendency to be a correct solution than if you didn't use an overtly positive prefix, even if that leads to a lot of annoying behavior.
- AlecSchueler 11mo agoI've always been taught to slightly overestimate how long something will take so that it reflects better on the team when it's delivered ahead of schedule. There's bound to be a bunch of similar advice and patterns in the training data.