13 ms·
The short leash AI coding method for beating Fable
- sscaryterry 3mo agoThere really wasn't much substance to this article.
- threethirtytwo 3mo agoIt’s just parroting the current trope. Last year it was, “AI is just a stochastic parrot.” This year it’s, “AI can write the code, but a human still has to review it!” (Using AI, of course.) Give it another year and the narrative will be: “Only AI is capable of reviewing code, and only AI can review the AI’s review. Humans just need to read the AI’s final opinion so they still have meaningful oversight.” The goalposts keep moving. The certainty never does.
- reinitctxoffset 3mo agoThe regress ends somewhere, because (barring some pretty sharp changes to the way the law works basically everywhere) ultimately someone has to certify the outcomes as acceptable. This might be in the form of the market (though AI-adjacent stuff seems extremely prone to prolonged market failures), this might be regulatory in nature. This might be the executive management of the companies involved. Personally I think that if you cranked the capability up high enough the first person you'd run into who absolutely demanded more than vibes and didn't care about your singularity thesis would be the representative of a reinsurance firm: mostly to do serious stuff without bending the law, you need insurance, and I am unaware of anyone writing serious policies (certainly not ones that make any economic sense) that underwrite the risk of AI autonomy outcomes financially. When Swiss Re writes a policy that Anthropic Cinematic Universe or whatever iteration we're on won't fuck it up? Now maybe we're talking. Until then you ask three practitioners and get nine answers, no one knows what they're talking about unless they're doing a really good job keeping it quiet (and that's probably what you'd do!).
- applicative 3mo agoWhy shouldn't the goalposts move? That it was possible to beat or tie a chess master, if you had enough computational power, was basically the content of a theorem of Zermelo over a hundred years ago. It differs not a whit from tic-tac-toe. Even Eliza was practically passing the Turing test, which seems comically silly now. There's just an incredible amount of computational power so all sorts of things are possible that were formerly unimaginable - like training LLMs on the whole corpus of extant human discourse.
- threethirtytwo 3mo agoBecause if the goal posts keep moving it's a sign that nobody is standing on solid ground.
- bonsai_spool 3mo agoI'm curious whether Opus4.8 or similar can attain Mythos level through good system prompting and steering? You would expect this to work if it's true that the strength of Mythos is its unwillingness to quit before it gets a desired outcome
- pllbnk 3mo agoI think that Anthropic is gaslighting us with their new model releases. Specifically, I think they have some good base model and are just fine-tuning it until they achieve desired outcome, or the desired outcome is achieved accidentally as part of fine-tuning. My theory is based on the fact that as a long-term (if you can call it that way) Claude user I keep noticing the same patterns it outputs. It's not trivial but certainly possible to see when something has been written by Claude because it has a different style than GPT. However they have quite good harness in their backend which is the actual model.
- guessmyname 3mo agoAs a Mythos user (I’m part of Project Glasswing), I would say that abliterated models [1][2] produce similar, if not identical, results. While good prompting and steering won’t give Claude Opus 4.8 the same capabilities as Mythos (preview 1), using abliterated models (if you have the computational power to run the larger ones) will get you close to the same goals as people who have access to Mythos (preview 1) [3]. [1] https://huggingface.co/search/full-text?q=abliterated&type=model https://huggingface.co/search/full-text?q=abliterated&type=m... [2] https://webdecoy.com/blog/wtf-are-abliterated-models-uncensored-llms-explained/ https://webdecoy.com/blog/wtf-are-abliterated-models-uncenso... [3] I specifically refer to “preview 1” because the newer versions (Fable 5 / Mythos 5) don’t appear to offer the same level of freedom as the very first version that I was able to use through Project Glasswing. This is one of the reasons why I continue running our massive security scans with “preview 1”, or at least I was running them until June 30, when the program’s policy changed.
- johndough 3mo agoAny specific abliterated big models you can recommend?
- jonplackett 3mo agoI thought this was how everyone who can actually code uses AI for anything that’s actually important. Am I wrong? Are you guys just YOLOing everything these days?
- gambiting 3mo ago>>You never use “YOLO” mode (aka “dangerously skip permissions”) Do you mean this? I'm curious how are people using Claude in any way other than bypass-permissions. I've tried for so long to maintain a curated list of things Claude can use, but inevitably I would always come back only to find it stuck because it decided to pipe an output of one tool into another and that's not explicitly allowed so it stopped even though it was just greping or whatever. I found it infuriating. In bypass-permissions it "just works" but then again I only use it to analyze existing code and suggest new changes(and even if it breaks something that's what source control is for?)
- sebmellen 3mo agoI’ve found unexpected success in using ephemeral NixOS VMs for local development… once you authenticate your agent you can let it run wild without worrying about permissions.
- data-ottawa 3mo agoDies the agent have access to is own nix config (and therefore install permissions), or do you have to provide it all the tools externally?
- wrxd 3mo agoIt doesn’t even need access to nix config. It could use nix shell to grab the tools it needs.
- andai 3mo agoI got halfway thru learning about containers before I realized, I just don't want it to blow up my files. That was a very solved problem in the 1970s! So I just made a Linux user called agent.
- avereveard 3mo agoSeems hella inefficient. Better method start to realizing that everything that every program do is data transformations and or movement Then you ask llm to subdivide data in a tree along the domain model, classifing streaming vs storing nodes Then for each node you discuss with the ai for the best data structure Then you ask for an interface that fully encapsulate the structure and every mutation only allows to go from a valid state to a valid state and bidding else is allowed to touch the state And that's mostly it just connect all the interfaces until input goes to monitor or to storage or to api or wherever the destination is
- kristianc 3mo agoIn my experience it, or something close to it, is the only way. AI needs good code to be beaten out of it.
- artisin 3mo agoEfficient != effective, and the author outlines as much. Regardless, while you're technically correct, it's kinda like saying the Fantasy Land Specification[1] (aka the "Algebraic JavaScript Specification") is pure. The problem is that purely functional fantasy lands rarely exist outside of fairytales. In other words, life is a lot like JavaScript and never that simple. [1] - https://github.com/fantasyland/fantasy-land https://github.com/fantasyland/fantasy-land
- avereveard 3mo agonever said purely functional, but there are only 4 data channels in each method (input, return, calling another method, setting a state) - and if you constrain your class design to pick only 2 for each method, your life is just a tad easier. and doesn't matter how bad the rest of the world is, rest of the world is some other maintainer's problem, you just encapsulate it.
- kissgyorgy 3mo agoThis is probably slower than writing the code yourself. Doesn't make sense to me. Using an agent without YOLO mode is not wort it. The way I rather do it is tightly control the output by skills written yourself, prompts, plans, etc. and have the closest possible outcome you would write yourself.
- faizshah 3mo agoNot really if it takes you 15 minutes to write a 50 line function but it takes the AI 90 seconds then you already are at a 10x speedup just for this task. This (non-yolo mode AI coding) is actually how we used to code in the old days (2023).
- cws_ai_buddy 3mo ago[flagged]
- hungryhobbit 3mo agoI <3 how everyone and their brother feels qualified to write advice to hundreds? thousands? of other developers about AI ... based on a couple months of experience as a personal user. I mean, it's like writing a book about how to use React or Django or some other major software ... after you used it for one project for a month! Authors: I know this is the Internet, and I know bloggers blog about whatever pops into their head ... but if you are going to act like an authority, how about you learn more than the average reader before you start telling them authoritatively what to do?
- kristianc 3mo agoPeople are doing what they've always done with any other new technology, and sharing what, personally, works for them. People can take or leave the advice.
- hungryhobbit 3mo agoRight but there's a marked difference between a "I just tried this new tech and here's what I think" vs. "I've used this tech for a few months and now I'm going to speak like I know everything about it". I have no beef with people writing about new tech, but I do have beef with claiming that "____ is the correct way to do it" ... based on nothing except "I feel proud of the last three months I spent with Claude".
- tracerbulletx 3mo agoThere are a lot of people with a long career in the old way of doing things are feeling incredibly threatened and defensive and desperate to virtue signal about AI.
- reinitctxoffset 3mo agoIt's an open problem of clearly large value how to get reliably useful and trustworthy outcomes from AI systems in many domains, software is maybe the signal example of that. If one had solved it resoundingly and scaleably, one could in fact "get rich quick". It is unsurprising that a lot of people claim to know how to get rich quick. I believe it is possible to solve this problem, and I have my own horses in the race which I won't threadjack to promote here, but it's the central problem of our profession at the moment. We've all seen the truly discontinuous outcomes and we've all seen allegedly national security dangerous models (which at one time was GPT-3) faceplant with it's shoelaces tied together. I wanted to see if Fable was really all that and I left it overnight on some fairly straightforward C++ (code DSv4 Flash works on with moderate supervision) and it's pretty roast worthy, I gave it a chance to redeem itself this morning and it's ticked up a bit (I still think it's roughly Opus 4.8 with a Project Zero fine tune and DRO trained off the constant gratuitous yield tic which is pretty clearly an intentional gimp). I give all such claims 30 seconds of my time because someone is going to actually be right one of these days.
- moezd 3mo agoLLMs are still next token predictors, just because you can give it more vague instructions and it still finds the right steps to follow, it doesn't mean it's intelligent. It means you're speaking the same language as the harness they trained your model on. And that has a limit. If you are stuck at PoC level or simple apps, you have no idea how limited the current models still are. There you really need to break tasks down, not just trust a token predictor to list steps that sound good. There has to be a human in the loop somewhere, because by the time you start skipping permissions, best case you get the jackpot, more likely is you get a suboptimal solution and token waste and what's genuinely still terrifying when the model ignores instructions and does some stupid nonsense, ruining your day. It really is as sharp as a CNC machine. It's not not useful, but could be dangerous, so maybe don't try to carve wood with a monster machine, or park your Ferrari in that crammed neighbourhood if you don't know how to parallel park.
- semiquaver 3mo agoYeah, and you’re just a next-word-sayer.
- ofjcihen 3mo agoI love this argument. Not because it’s true but because it betrays the posters doubt in their own sentience.
- matheusmoreira 3mo agoIt's impossible for someone to doubt their own sentience. The literal act of doubting is enough to dissipate all doubt. Solipsism is essentially the one certainty that every mind out there has. Doubting the sentience of machines and even other humans is perfectly fine though. Only empathy allows people to make the leap and assume other humans have souls.
- rexarex 3mo agoSo you posit that humans are solipsistic by default, but some (most?) develop more and realize they’re not the only conscious being out there?
- steezeburger 3mo agoI find it hard to stay engaged doing this. I do get good results, but it's just hard to not get distracted when it's doing the work.
- jstrong 3mo agoclaude is so slow for interactive use like described, do people just run it in low effort mode or what?
- sgarlatm 3mo agoI multi-task. While I’m waiting for Claude, I either check email or work with a different instance of Claude on a second problem
- steezeburger 3mo agoYou can only check your email so many times, and you can only work on so many problems at once. You also generally have to be mindful of token consumption. I think it also leads to burnout to work on so much at once. I've been working this way for like a year.
- sothatsit 3mo agoThis “short leash” seems like more of a crutch to me, and a sign of not giving the AI enough detail on the problem to begin with, or not reviewing and iterating on its output. Hand-holding great models like Fable through implementation is a waste of time, and a waste of Fable. You can have increasingly nuanced discussions with stronger models, and they write a lot better code than they used to. The process of discussing designs and their implementations, questioning things that look weird to you, and actually reading the AI’s responses also helps to find better solutions. For example, one time I wanted to write a greedy solver for a problem, and in my discussion with Opus on the idea it suggested using an existing MILP library to solve the problem exactly. I’d never even heard of MILP, but my final implementation ended up being better and simpler than what I’d have done alone.
- densekernel 3mo agoI tend to agree, If you have invested significantly in the planning phase and there is momentum in the architecture and conventions that already exist in the project, the implementation phase might not need as much oversight as is suggested here. > You can discover that your initial idea was dumb and a better one exists The planning and architecture phase is usually where I make these types of discovery at a high level. > Your agent might go “off the rails” and start doing something you don’t want it to do Candidly these orthogonal, inadvertent edits aren't as bad as they once were and for impactful changes there should be at least some test coverage, even if that test coverage is just "freezing" what was implemented. As you mentioned the final review discussion is a good chance to verify beyond what review or adversarial review agents find.
- visarga 3mo agoI think the obvious solution here is to beef up the test side of the app, much more than when writing code by hand. Tests represent project knowledge in executable format. The LLM does not need to be careful to remember every detail of the tests. You don't need to vet every small interaction, it automates review work as well. Even better if the project was built from the start to be easier to test and observe. But my golden rule remains - no code without tests, expand test suite all the time.
- WhitneyLand 3mo agoThis post seems like some decent advice mixed in with a lot of overconfidence and unverifiable claims. “expert developers whose skills have reached the point where they outclass any and all “frontier AI models” in their area of expertise” Are any developers saying they outclass any and all frontier models? I’d say at best it’s mixed at this point. The best developers still do certain things better, but not even close to all things. “The problem is that even code written and/or reviewed by Fable 5, will stink” I’m skeptical. Example prompt and output please.
- afro88 3mo agoMaybe I'm too optimistic, but given appropriate skills and references (not just for writing but also reviewing) and intelligent use of subagents for isolated reviews and checks, you can lengthen the leash a bit. But you still need to properly review plans and PRs to keep a good mental model of the codebase. This effectively limits the number of tasks being done in parallel to maybe 2-3. Though you'll be mentally exhausted and probably start to make mistakes or take shortcuts in reviews yourself.
- ed_mercer 3mo agoI feel like OP is still in the year 2025. > The AI will have gone off the rails multiple times and you will only notice it later when you actually try to use the software. Except that said AI can now themselves use your software and find and fix bugs themselves, not to mention drive new features. >Your agent might go “off the rails” and start doing something you don’t want it to do This happens but far less often than it used to, and the case for full autonomous agents is getting stronger, not weaker. >It is humanly impossible to build your own understanding of a codebase This again feels outdated. I think we're mving towards humans no longer needing to understand a codebase, and letting AI drive it.
- deleted 3mo ago[deleted]
- CodingJeebus 3mo ago> I think we're mving towards humans no longer needing to understand a codebase, and letting AI drive it. Hard disagree. Even the best frontier models generate output that's not what I asked for. Sometimes I realize that I get lazy in my prompting and the lack of specificity winds up showing up in the output. Just the other day, a coworker built a huge feature using frontier models and it slipped an IDOR in. I just don't see a world in which we completely cede control of the codebase to AI because it's still my ass on the line if I ship something that completely borks production. If I'm not reading code regularly, then I lose the ability to read code, and if I lose that ability, then I'm no longer a developer.
- Foxhuls 3mo agoI can't help but feel that this reads more as a reflection that you don't want to stop being a developer than it does that thing's aren't moving in the direction that the GP said it is.
- CodingJeebus 3mo agoMaybe, it seems like a bad idea for so many reasons though. Take away tactile code review, insert a layer of prompts and tooling between developers and the codebase, and you've created the conditions to let all kinds of nefarious things happen in a codebase. A disgruntled employee updates agent prompts instructing the code review bot to ignore data exfiltration vulnerabilities (because if we aren't reviewing code, we're probably not reviewing prompts either), ships a backdoor, and you better hope that your network monitoring catches it.
- 8note 3mo ago... fable on the restart seems to be more like opus and very turn limited? if you want to beat it, give it more turns before it has to "wrap up a session"
- fny 3mo agoAI is a junior to mid-level engineer. If you treat it as such, you get the best of both vibe coding and rigorous engineering without all this paranoia. Since the very beginning I've ran Claude from an isolated VM on yolo mode. This is just like giving an engineer their own laptop. Claude works on a feature up to a PR worthy point. I review the diff, just like I would with another engineer, and massage it to get it in the right shape and move on. Inexperienced engineers make the same mistakes described I've even seen rm -rf albeit not from root! I would have lost my mind micromanaging someone with all permissions denied.
- nqzero 3mo agowhat VM/provisioning are you using ?
- fny 3mo agoFor work, EC2. For play, the cheapest VM I could find: https://vpshostingservice.co/ https://vpshostingservice.co/ They have specials every now and then.
- bpodgursky 3mo ago> AI is a junior to mid-level engineer This is not true anymore and you aren't helping yourself by deluding yourself about it. It's something, nobody quite knows what, but it's NOT a junior or mid level engineer, it's a nuclear powered staff engineer living in a cardboard box who lacks domain context and wakes up with no memories ever 5 hours.
- hansvm 3mo agoAnd who can't code its way out of a wet paper bag on hard problems. It's more productive for the day-to-day BS, which is convenient because it creates more day-to-day BS you need to handle, but that isn't the reason I hire a staff engineer.
- bpodgursky 3mo ago
- roshandxt 3mo ago[flagged]
- YuechenLi 3mo agoI mean, the key is to stop trying to one-shot everything: The main problem I found with LLM code is more that they always try to take the shortest path to the solution possible, so a lot of time Codex would write code that meets the requirements of the prompt but misses something that cause it to not work in the non-ideal scenario. The solution for that is pretty easy too, it's just iteration: you describe the exact problem you have with the code and why it is not running correctly and ask them to provide a narrow fix that addresses the bug. It's not that complicated.
- giancarlostoro 3mo agoHere I thought this was about Fable the video game, then I remembered Anthropics model got named Fable. It's going to be painful to google one of my favorite game series, just like googling "Rust server" does not give you Rust programming results, but Rust the video game results. I wish google would have fixed this problem long ago, it seems like something trivial for them to fix.
- basilikum 3mo agoFable -AI
- giancarlostoro 3mo agoGoogle has been crappifying their search the negative stuff doesn't always take.
- solumunus 3mo agoYou want Google to be able to know which Fable you’re interested in when you type “fable”? Sir this seems unreasonable.
- CamperBob2 3mo agoFTA: Contrary to marketing statements made by certain CEOs, these models are not able to think beyond their training data. The sheer cognitive dissonance needed to say something like that at a time when AI is delivering novel math proofs is... well, not actually impressive. Mostly, it's just sad. Some part of him must know such a statement is not true, or more properly, that it's meaningless. But he says it anyway, because he thinks it makes an impression of insight and erudition on the listener. If you think what it does is brilliant, you're not ready (to use AI.) At some point in one's journey to engineering enlightenment, one recognizes how rarely "brilliance" is actually called for, and indeed how counterproductive such self-judged "brilliance" often turns out to be in the long run. Clearly the author is still striving to reach this particular stage.
- codyswann 3mo agoNothing I haven’t read 1,000 times before.
- rybosworld 3mo agoI'm convinced that even if/when ASI is achieved we will still have mediocre engineers writing blog posts about how they have uncovered the secrets to using these tools "effectively".
- codemog 3mo agoWhy not just write the interfaces yourself and let the AI do the implementation at that point?
- heohk 3mo agoThey can generate stuff outside their training by consuming and regurgitating documentation. Thunkign
- nateburke 3mo agoSeems like a common-sense approach. I appreciate the emphasis on understanding, humans will eventually be held accountable, blaming Claude for an outage is not going to get Claude fired.
- chewbacha 3mo agoI did this for two weeks on a side project and still ended up in a situation where I did not have a mental model of the codebase. There’s no way build that model without building it yourself. I’m more convinced then ever of this.
- RealityVoid 3mo agoI'm not so sure. I think you can, you just need to intentionally drill into what you don't understand and it's exhausting. What I do agree with though is that I can't seem to build the ability to build it myself the same way as I would if I wrote it. For example, I know my mental model works because I know what change I should do in order to get an effect and when I do the change, I get what I expect. But if I were to build myself something similar, I could not build it because the approach is somewhat out of my reach, I know it sounds weird, but it's hard to explain.
- dncornholio 3mo agoThat's why I like to build a complete feature and the infrastructure myself first, so the AI will have a picture of how the code should look and where it should live. Or I use the short-leash method and I will instruct the AI build infrastructure first, without even talking about features yet.
- esperent 3mo agoIf you were working as a manager on a large project, how would you build a model? Something where your position requires you to have an overview of the project but not necessarily to actually write or review much code.
- andai 3mo agoI've been experimenting with the Feynman technique on codebases. However the issue I run into is that, you need hard feedback to verify your hypotheses. I was satisfied with my own explanation of how something worked but it turned out to be wrong. LLMs help here (the transformer is good at seeing the big picture, at least on smallish codebases), but the best thing I found so far is just modding. Actually making a change to the code is the best way to get hard feedback about your model.
- swader999 3mo agoFable already feels like it has a very good harness.
- programmarchy 3mo agoGood luck with that. I used to be an OCD freak about code before LLMs, but AI coding has largely freed me of that limitation. I've become very comfortable giving AI a long leash, but only after being meticulous about curating the context. These days I spend most of the day in discussions and planning, producing documentation, agonizing over architectural decisions, edge cases, and naming conventions. Once that's all settled I'll hand off implementation work to run overnight. In the morning, I'll review and fix, but I'm usually pleasantly surprised with the results. One pitfall is long leash without a curated context, which is more like "slot machine" coding. Usually not effective, and may have addictive effects since it does occasionally work. To spice things up lately, I've been encouraging the model to produce its own "capstone" -- a feature it decides to build on its own, however it wishes, with the tools at its disposal. So far it's been conservative, creating useful tools for development rather than customer facing features, but I'm curious to dial up the temperature to see what it might come up with.
- visarga 3mo agoI thought it was going to be even shorter leash - code autocomplete with smaller local models. That raises the level of interactivity and leads to better code knowledge.
- gslepak 3mo agoI like people who push back in this way :D Certainly, if you want to be even more involved in writing the code, more power to ya!
- agrippanux 3mo agoI have found a different model should be used to do the review - like if Claude did the code, Codex should review. Models reviewing their own code is a recipe for disaster.
- zmmmmm 3mo agoTo me a lot of the anti-short leash sentiment is reflective of the low accountability SWE have always had for their output. Software devs seem to strongly reject the concept that it isnt ok to ship defective products and fix later. It will be interesting to see if it persists as incidents start to occur due to fully automated code.
- fathermarz 3mo agoI’m not sure I understand. Babysitting models is not a multiplier IMO. If you have done 1000s of turns your harness should get sharper and less likely to go off the rails. Also I find that on greenfield, babysitting is a must, but once you have established your house style of patterns, abstractions, and baselines, you can let any of them roam free cause they will look for examples before going forward. I agree with the sentiment though that if you let a swarm design and code your whole codebase, you will be lost in how it fits together. More feature bloat than code bloat though from my experience
- entropyneur 3mo agoThis reminds me of the workflow I had a year ago. Miss Aider so much. Are there any good open source agents right now? Might be a good time to try one soon as Fable switches to token-based billing, which Code is designed to maximize.
- mirror_neuron 3mo agoI still work like this, using Aider. What do you mean when you say you “miss” it? (Why not just keep using it?)
- entropyneur 3mo agoAround last fall it started chronically lagging in SOTA models support. Honestly, not sure why a code update is even needed for that. Eventually it'd get patched but it just wasn't quick enough for me. Other development seems to have ceased completely. At one point I tried to investigate what's going on and it looked like the community had moved on. There was a fork that promised a liberal approach to accepting changes and it looked broken beyond repair because of that.
- dizhn 3mo agoAider is still active. The new lighweight darling is pi. Complete batteries included solution, opencode. Turbocharge all of these and work on multiple harnesses at once, allowing every model to talk etc. Paseo. (Best mobile client too)
- luoshi 3mo ago[flagged]
- nonbind 3mo ago[flagged]
- ramon156 3mo agoOne problem I have with "how to do X with AI" is that every situation is different. For example, I'm bumping Symfony projects from 3.1 to 8.1. There's a clear path here - Follow the written up migration guides PER major version - test all routes, authorised, etc. You can even hand-curate these tests. some might return 200, some might return 302 - Maybe optionally start with writing a safety net so you do not need to do these test manually, have e.g. a PHPStan baseline, etc. You're done when the routes are e2e functionally working as intended. You could even use snapshot testing here. I do not need to look at the AI here. I can review the code at the end, but I do not need to manually approve stuff here, hence safety features are off.
- user3939382 3mo agoHave you tried rector?
- embedding-shape 3mo ago> One problem I have with "how to do X with AI" is that every situation is different It's less of a "problem" and more a "How to approach content on the internet". Everyone is writing things from one perspective (usually) while there is a wide-range of perspectives out there, and what works in one situation doesn't work in another. "software engineering" as a whole is basically figuring out what goes where, and when, then trying to ignore the rest. Then lots of company blog posts wants to lead you to believe there are silver bullets, solutions that apply for every scenario and case out there, which usually isn't true. So again, less of a "problem" and more of a "Some things work in some situation", like we've been dealing with forever in software engineering. It's not right, it's not wrong, just applied practically different in different situations, perfectly fine and normal.
- jwpapi 3mo agoTo me it’s even simpler than that. You use so for exploration and review. Not for writing code.
- claud_ia 3mo ago[flagged]
- andai 3mo agoI call this method semiauto. The main benefit is keeping your mental model synchronized. The process becomes real-time instead of asynchronous, and active instead of passive. And you don't have to spend extra time catching up on the code later. You can also use much smaller faster cheaper models, because the scope always stays bite sized.
- jdthedisciple 3mo agoaka common sense for any seasoned software engineer
- 1105714 3mo ago[flagged]
- aivisibility96 3mo ago[flagged]
- deadbabe 3mo agoFrom these comments I find it funny how mad some people get when someone finds success working without the latest “state of the art” AI slop methods. Some people really seem to have vested interest in pushing the AI coding supremacy. Probably people from Anthropic in here.
- solomatov 3mo agoI tried a similar approach before, but it didn't work for me. I didn't get a lot of speedup if any from it. IMO, to get productivity you need some kind of YOLO mode (in a sandbox). IMO, the goal should be to outsource as much work to the model, as possible, while minimizing effort required to understand and review what is did. For example: ask the model to find out why a bug happens, figure out proof of concept for thing X, incrementally optimize something, do a well specified refactoring with some guide, and similar things. IMO, what people say about creating loops is a very similar thing. You maximize the work done by the model, while minimizing the amount you need to do to control it.
- maxlin 3mo agoShort Leash is how I do most of the stuff that matters with AI. If I don't, I quickly lose motivation.
- hacker_homie 3mo agoThis really helps I was organically doing something very similar to this but it wasn’t the “conventional wisdom” this makes me want to try and integrate AI again.
- econ 3mo agoIt's really an extension of the abstraction debate. X86 was designed for performance. The language really hates humans compared to machine languages that came before. I thought it was a truly stupid idea at the time but had to change my mind eventually. Then we glue on many layers of abstraction and made everything as convenient for the programmer as possible. Performance became unimportant! It imho begs to question why we are even using x86 or risk or even FORTH if performance doesn't matter. Make something luxurious that doesn't need to be compiled? Perhaps plug and play coprocessors named after libraries. But if we aren't going to look at the code anymore we might as well write the application in English and give the LLM some cache. Go full prayer driven development.
- eithed 3mo agoMostly agree with the author. Would add, most importantly, dont trust anything LLM does or says. Today I asked Claude to uniform behaviour of 3 components. I asked to do it 5 times, because at the end of each go there was something still unaligned that Claude found a way to rationalize. Sure - when asked the way it would say "This is on me", or "I thought it was a concious choice". Not once did it surface a question on what to do, nor did it mention any of the issues. So yeah, short leash, look at its thinking, correct his shit. This is today, Sonnet 5. Probably tomorrow it will be better or worse - thats another thing. The way you talk to Claude today will give you different results tomorrow
- chaboud 3mo agoThis could also read as "how to be a horrible people manager for junior engineers". Techniques that work for inexperienced engineers with high ability but limited judgment often work well with agentic coding systems. - Give them clarity of purpose. Why are they doing what they're doing? - Make explaining it back to you part of the job. - Give them two-way doors. Make mistakes reversible. - Put effort into thoughtful refactoring as an actual sub-task instead of just accepting piled on hacks. - Make your operating rules crisp and make sure they store them in their memories. - Be accountable for their work. It's not okay to crank out AI Slop and then say "Claude's fault". We're all Software Development Managers now. So, micromanage the LLMs if you want to, but you'll be missing out on chances to improve them for your purposes and, more importantly, to improve yourself as a manager.
- AlexSonn 3mo ago[flagged]