3 ms·
Great site, triggered memories! haha. To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen
by dudeinhawaii 27d ago
Great site, triggered memories! haha.
To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful".
Models nowadays want to double-triple-quadruple check things. I'm being silly but it verges on "I have a working solution but let me write a variation in Rust to ensure a convergent solution and prove this works".
I've had to stop models nowadays mostly because they're being agonizingly pedantic in their validation. Opus is actually one of the most pedantic and "off track" here. But again, not in a bad way. I'm usually like "stop testing latency between 50 runs of this app... this is version one.. we're going to make a million more changes.. you're not buying us anything".
- mrinterweb 27d ago> triggered memories Yeah of 10 minutes ago. It is shocking how long some seemingly simple things can take. I know there are some things I can do faster than the LLM and some things it can do faster than me. The amount of rambling BS is the exhausting part.
- epistasis 27d agoTry different models, it's a breath of fresh air. GLM 5.2, etc. all make life much more enjoyable. They may not one-shot a complex project the same way that Claude can spit out memorized architectures, but that sort of system is always only useful for a one-off prototype anyway, so not much is lost. Edit: for how to do this, I set up an OpenRouter account so that I could easily switch models, and then ran them in Pi inside of Orca ADE. Orca lets me easily switch from Pi to Claude Code to Kilo to Codex or Hermes or whatever. Pi+OpenRouter lets me easily switch the LLM. All of it lives in a single open source orchestrator to avoid any platform lock in to any AI company going forward, and can even do local LLM should I care to take on that massive hassle.
- torginus 26d ago> Yeah of 10 minutes ago. It is shocking how long some seemingly simple things can take. And shocking how little code end effort some things take if done by hand. I am not some hardline LLM hater, just venting my frustration.
- sixsevenrot 26d agoDoes it help if thinking is not set to High?
- f055 27d agoClaude models were always too eager and "overly helpful" for my taste. But it seems better models tend to be this way. GPT 6 and 5.6 are overly helpful too, but at least less than Fable. But I seem to be sticking to GPT 5.5 as this was a really focused model.
- darepublic 27d agoI blame the hidden context on the tools/subagents. One recent example.. codex can just look in the code for Db schema but continually tries to request permission for a live db query
- qurren 27d agoOne thing it's missing: "smoking guns" and "smoke tests" If you search my company's Slack for "smoke" the results are almost all within the past 2 years ...
- digitaltrees 26d agoWhat, you didn’t smoke test before AI? I mean did you even really code then? :)
- testplzignore 26d agoThis is the load-bearing question.
- qurren 26d agoAnd it deserves an honest answer.
- bot403 26d agoThe answer is genuinely yours
- digitaltrees 26d agoAnd that changes the game. Again.
- 26d ago
- SamuelAdams 26d agoThis is my recent experience as well. Models want to run linters, tests, etc. And that is all covered in GitHub actions. So I have been instructing agents to push a draft PR, then I validate the static checks pass and tell the agent if there are issues. Agents and AI are getting expensive, it seems silly to waste tokens on static checks.
- CharlieDigital 26d agoDownside would seem to be that CI tends to increase the cycle time and feedback loop and add their own cost into the equation.
- mejutoco 26d agoIt seems running those tests, linters, etc locally would save pipeline minutes as well, like precommit hooks
- bahbahbahbah 26d agoPlus rigorously ensuring backwards compatibility for a project that is 2 hours old and has zero users.
- cruffle_duffle 26d ago"Plus rigorously ensuring backwards compatibility for a project that is 2 hours old and has zero users." That is exactly how the slop accretes and you get a pile of crap. Claude somehow assumes that said 2 hour old userless app is some dusty enterprise app with millions of users and billions of dollars at stake for a 1 second outage. I have to constantly have these things "take a deep breath, step back and look at the entire thing and do this change holistically. please restate what i'm asking you to do and why it's important"
- butlike 26d ago> please restate what i'm asking you to do and why it's important" This doesn't actually do anything though, right? There's no understanding, so the machine will just reiterate the original token query back to you. The 'why it's important' part will just generate some patronizing boilerplate as a raison d'etre.
- Filligree 26d agoThe purpose of a prompt like that is to have it rephrase it in its own words, so you can spot any misunderstandings and correct them. If you're of the belief that this is a contradiction in terms, then what can I say? It works.
- johannes1234321 26d agoBut it doesn't "understand" these words any better. When transforming into code it may reinterpret yet again.
- Dylan16807 26d ago
- yonatan8070 26d agoI've noticed them repeatedly casting the same value to the same type for now reason, like I'd have a Python function with a type-annotated int argument, and inside the function it would cast that int to int, and also at the call site, just in case it wasn't int enough.
- ls612 26d agoThis is mostly a side effect of post-training models to not hallucinate, which has obviously been a major priority for a while now. They are highly incentivized to double check things to avoid accidentally making stuff up.
- classified 26d agoWell, they have to crank up the token spend somehow!
- koyote 26d agoIt basically makes it impossible to have it build a very large feature incrementally. I am currently building a very large feature and the only way to do this properly is to do it in small steps. No matter what I prompt, I always have to spend a bunch of time reigning it in during planning (and implementation!) because it tries to solve the whole thing end-to-end or add enough boilerplate so that a code path (that is still under construction) returns a meaningful (but wrong) value. And then of course every time I start a new session it gets very upset that the feature I am trying to implement will never work because the main user entry point has a 'not fully implemented' exception that definitely needs to be removed immediately!
- jchook 26d agoI typically see the exact opposite problem when giving Claude the reins on a large undertaking. Claude will, without prompting, break the implementation into 6 phases, and write AI slop “code as English” specs for each phase, each one with glaring errors and unintelligible terse jargon. It will review them all several times with major findings every time and tons of design churn… Then it will implement a total heap of garbage over many hours of many agents, despite it working in “lanes” and in parallel, and with regular input needed. Just thousands and thousands of lines of junk, which auto review then plays whack-a-mole to polish and fix. All the while it is able to see the errors and edge cases, yet fails to see the key architectural blunders that led to them in the first place. I’ve had to fully rethink my approach to LLM tasks like this. For example, for library-esque modules, I have found that isolating the problem outside of the codebase is one useful approach. Something about the lack of noise. It can land on a cleaner solution that can be retrofitted. I also find that asking it to implement end to end in one fell swoop with a very high level plan actually saves time and creates a cleaner result.
- genodethrowaway 26d ago"triggered memories" sums it up, yeah. Haven't seen behavior like this in a year? year and a half, maybe? Lord, this stuff is moving fast. Hope we figure out alignment