7 ms·
Agentic coding notes
- brcmthrowaway 3mo agoThis seems like the beginnings of AI psychosis, tbh.
- deleted 3mo ago[deleted]
- mock-possum 3mo agoHow do you figure?
- MomsAVoxell 3mo agoThe immensely desperate search for meaning in the mundane, perhaps?
- zuzululu 3mo ago[dead]
- zarzavat 3mo agoFable changes the game yet again, because it's API-only. You're not likely to want to run Fable in a loop any more than you want to take a bunch of dollar bills and light them on fire. Every invocation of Fable has to be intentional, its context carefully managed. I feel like a babysitter.
- eru 3mo agoFor now, I can use Fable from the web just fine. > You're not likely to want to run Fable in a loop any more than you want to take a bunch of dollar bills and light them on fire. Every invocation of Fable has to be intentional, its context carefully managed. Eh, that's just because it's the current frontier model. Give it a few weeks, and prices will drop.
- zarzavat 3mo agoAPI prices are the new normal. I doubt that prices will drop to the level of the subsidized subscriptions any time soon. Usage is growing exponentially but capacity cannot. There is no reason for them to waste their capacity on subscription users if they can sell that same capacity to API users. Like with Uber and Lyft, the low prices were a fight for market share, but now they have successfully captured that market share the focus changes to balancing their books.
- eru 3mo agoI suspect subscriptions will stay. But you will see more and more roadblocks that the 'barely goes to the gym' user barely notices, but that the power user will chafe under. I make that prediction, because the people who pay for subscriptions but only use them moderately at best are truly profitable.
- baq 3mo agoThe $200 sub is the new free tier, has been for a while now.
- weird-eye-issue 3mo agoCompared to Opus 4.8 I really haven't been impressed
- danielbln 3mo agoAnd I've been quite impressed. Opus talks the talk, Fable walks the walk.
- stingraycharles 3mo agoI don’t understand what these comments add to the discussion, you always see these and it’s just noise at this point.
- danielbln 3mo agoThey add nothing, meaningless anecdotes. I was kind of riffing on that.
- skrebbel 3mo agoPosting meaningless comments isn’t suddenly useful because you’re aware of it and “riffing on that”.
- danielbln 3mo agoAnd now you've added to the pile, congratulations.
- weird-eye-issue 3mo agoSo you added BS, on purpose? There is nothing meaningless about anecdotes. The plural is data
- MomsAVoxell 3mo agoYeah come on, if these endless dialectic statements get made, at least post more detailed information about how your position was attained.
- NitpickLawyer 3mo agoI agree with you that you don't need fable for everything, and you have to be careful on what you run it on. CRUD stuff, sure even the small models can do it. But there certainly are tasks that are very much suited for the absolute SotA and you'd leave money on the table by not using it. And how much a task is worth is dependant on how much it improves your bottom line. So the cost/token becomes largely irrelevant. Let's take this [1] benchmark. A bit more context here [2]. Here models are asked to create kernels for running inference on models. This is a benchmark perfectly suited and highly relevant right now. It's easily verifiable, an active are of research, and the results are immediately useful. Say you have 1 unit of compute, it costs 300k $ and serves 1x users. In comes Fable and after one session it gives you 30% speed-up on your 1 unit of compute. It can now serve 1.3x users. How much is that one session worth for you? How much is it worth for a company using 10 units? 100 units? How much is it worth for a hyper-scaler running 10.000 units? How much is it worth for a lab that trains the next frontier model and then serves it from 100.000 units? 30% is relative. And the cost for one session is really meaningless. It can cost 1m$ / session and it would still be worth it for someone. [1] - https://kernelbench.com/mega https://kernelbench.com/mega [2] - https://x.com/elliotarledge/status/2072814573753975266 https://x.com/elliotarledge/status/2072814573753975266
- nobodycares1 3mo ago[flagged]
- jacobgold 3mo agoFable is supposed to return to subscription plans, unless I'm missing something: https://jacob.gold/posts/fable-5-removal-is-temporary/ https://jacob.gold/posts/fable-5-removal-is-temporary/ Anthropic says the change is about capacity and is temporary. In its launch announcement on June 9, 2026, it says: "After this point—when sufficient capacity allows us to do so—we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can."
- techpression 3mo agoThey can’t keep their current models working on subscriptions[1], so we’ll see if this is marketing or not in the future. It’s smart to tease it no matter what, ”insert classic first hit is free drug reference”. [1] https://status.claude.com https://status.claude.com
- zarzavat 3mo agoApple claims their price increases are temporary too. I'll believe it when I see it. The capacity limitations are not going away any time soon.
- vidarh 3mo agoI just had Fable run overnight in a loop, and it fixed ~150 compiler crashing bugs that Opus had kept deferring. I wouldn't start with Fable - when I use burndown loops I tend to include instructions to document progress and set aside anything that turns out to be harder than expected, and solve the easy stuff first. When a model runs out of easy stuff and start struggling to make progress on what is left, I can let it keep churning on that - they get there eventually - or I can bump it up to a smarter model if one is available. Opus had churned a week driving down spec failures, and did a great job. The 150 Fable took overnight were the ones Opus had kept putting aside.
- steveklabnik 3mo agoI had fable running in a loop overnight last night, finding bugs. It found a heap overflow. That triggered its safety guards, which converted the thing to opus, leaving opus to run the rest of the night, wasting my precious time with Fable. Oh well, it was pretty funny, all things considered.
- vidarh 3mo agoThat sucks - I've thankfully only had that happen once. Similar thing - I was testing my new X11 server, and it turns out that broken X11 packets can make Firefox crash, and I got a refusal on a sub-agent request Fable prompted to have an agent narrow down why.
- ianbutler 3mo agoThere's a setting that stops fallback to Opus 4.8 if you would like to avoid that behavior. Config -> Switch models when a message is flagged -> false That should stop it entirely when a message is flagged and then you can come back without Opus having potentially made a mess.
- nobodycares1 3mo ago[flagged]
- anon7725 3mo agoYou’re out of your element, Donny.
- bob1029 3mo agoA lot of the crazy ideas seem to have melted away in the face of massive context sizes. Today, I can put roughly a megabyte of utf8 text into my system prompt before things start to get weird. That is a massive amount of information even if we are being sloppy with it. You can read The Hobbit and the first Harry Potter book cover-to-cover and still have room to spare. I would deeply struggle to develop a world model this detailed for any business. Anything that needs to get more specific than these narratives can be a SQL query tool into the data warehouse, grep over the codebase, MS graph API lookup, etc. Giving the business a balanced way to collaborate over this one shared model of the world is a new challenge I am beginning to engage with. I've also noticed that the world model will compound on itself in terms of self-detection of update opportunities. The more constraints there are, the more likely we appear to violate one.
- stingraycharles 3mo agoYou’re forgetting that keeping the context small is better economically and delivers better results.
- themgt 3mo agoForgetting? I think you mean to say your advice was auto-compacted to keep our context small and deliver better results.
- bob1029 3mo agoI learned some nuance. Small context is important if the context isn't cachable. If most of it is the same prefix every time, the situation changes pretty significantly.
- embedding-shape 3mo ago> melted away in the face of massive context sizes If only. There is a huge difference between "Gives good responses/can easily spot things within N context size" and "Technically works but sucks within N context size", almost all models basically become cave-people once you go beyond 50% of the "supported" context size, meaning while they may technically work with 1 million output tokens, those last 500K tokens are gonna be massively "dumber" than the first 500k tokens.
- foobarbecue 3mo agoIt's "Galapagos" or "Galápagos," not "Galapogos."
- stingraycharles 3mo agoYou miswrote OP’s miswriting in the third version :)
- foobarbecue 3mo agoArgh, autocorrect got me. Thanks, fixed.
- foobarbecue 3mo agoAaand now OP has fixed it in the HN post title. Still wrong in the linked article.
- deleted 3mo ago[deleted]
- MomsAVoxell 3mo agoHe's not referring to the real islands, but rather the state of mind imposed upon existence on Vancouver Island.
- martey 3mo agoOP's alt text makes it clear that by "Galapagos Island" they mean Vancouver. I assumed that this was some sort of local nickname, but all of the references to "Galápagos of Canada" I could find are talking about Haida Gwaii instead.
- pluc 3mo agoNobody calls it that.
- skrebbel 3mo agoYeah I was like “woa a PGConf on the galapagos, i gotta get my ass to one if those!”
- progbits 3mo agoRealizing he's just using it to mean remote place in terms of AI bubble (Vancouver! What does that make all the other places that are not major tech hubs?) was a bummer. Who cares about AI, I wanted to read about living in Galapagos
- MomsAVoxell 3mo agoI'm saddened to have gotten this far and had my bubble burst, I legitimately thought someone was on the real GI's and was able to compose such a treatise, nevertheless.
- ballsac 3mo agoYou think humans consuming the Galápagos Islands is something worth aspiring to?
- MomsAVoxell 3mo agoIts more that I had hoped that someone with clear AI psychosis was in the right place for them to step back from that precipice and that we might all learn something valuable from observing it.
- gwern 3mo agoURL typo: "hange how he works](/productivity-velocity/)". (I make this kind of Markdown syntax error all the time and set up a lint for '](/'.) You should talk to https://www.mechanize.work/ https://www.mechanize.work/ for sponsorship/credits and about environments.
- deleted 3mo ago[deleted]
- duckmysick 3mo agoI'd like to highlight a different part of the article: > In general, when I talk to software folks about testing, I'm coming from such a different place that they immediately look at me like I'm an alien, so let's talk about how we tested at this hardware company I worked for, Centaur, which informs my biases about how I like to work. Some of the things that we did that were or are unorthodox in the software world are: > Hired dedicated QA / test engineers, with testing being a first-class career path on par with being a developer - No code review by default - Virtually no hand-written tests - Constant testing via what programmers sometimes called property based testing, randomized testing, fuzzing, etc., although we just called those tests (hand-written tests were called "hand tests"). - Large regeression test suite (3 months wall clock to execute on compute farm) - No unit tests Anybody here tried that (or a similar) approach? Especially going all-in on property based testing and fuzzing with no unit tests. I tried that approach somewhere before and the initial results were promising, but ran into political issues so the idea was canned.
- bitwize 3mo agoThe first thing I started wondering was "is this the same Centaur that comes up as 'CentaurHauls' on a CPUID (EAX=0)?"
- deleted 3mo ago[deleted]
- imrehg 3mo agoShould be, judging by their LinkedIn from the bottom of the page (Centaur Technology - Member of Technical Staff: 2005 - 2013) This is a blast from the past. Centaur was doing VIA Technology's CPUs, and I was at VIA (in Taiwan) while this author was in Centaur (in the US). I was on the embedded side, but I remember some distinct collaborations with the US team, so there's a non-zero chance to have crossed path.
- rtpg 3mo agoI really wonder what "randomixed testing" looks like in practice. What is the measure of success/failure? I undrestand for fuzzing you have a very basic "doesn't crash" metric. Property based tests.... you gotta write properties for the PBTs to work on. What is the randomized testing hitting?
- zapnuk 3mo agoThere is a reasone we use left and right margin/padding. This blog is quite unreadable for 27/32" monitors.
- Aldipower 3mo agoToggle to "reader view" or resize your window. It is up to you and really isn't that hard.
- zapnuk 3mo ago"You can lead a horse to water, but you can't make him drink" Reader view makes the text too narrow: https://imgur.com/a/yQqzxco https://imgur.com/a/yQqzxco Sure i could take extra steps to make it more readable, but at that point I rather not read it as im not that interested nor invested. Totally fine, no complains from my side. But If the author wants to increase the reach of their posts they just might use agentic coding to have the llm optimize the website for readability.
- Aldipower 3mo ago"Reader view" is adjustable via your personal settings, so I am not sure what you are talking about.
- layer8 3mo agoThere’s a reason not to maximize your browser windows. How do you handle HN threads on that monitor?
- baq 3mo agoWhy would you work with anything other than maximized windows unless you have a super ultra wide display?
- layer8 3mo agoBecause not all content is suitable to be stretched out 16:9 (or whatever the screen ratio of your desktop monitor is). In fact, in my experience most web site content isn’t suitable for it. The default browser window size I use is closer to 4:3 for that reason. There’s also software to auto-resize the window to different widths and to auto-place it to different horizontal positions depending on the website, that you can configure for frequently used websites.
- cognitiveinline 3mo agoIt's really coming down to "Do we want to subscribe to a human with a salary of many ~$10000(s) ", or "Do we spend 100(s)$ on an AI subscription" Even with it's issues, the latest models are going to disrupt the labor economics.
- layer8 3mo agoOn the other hand, you have many more humans to choose from than models, and they don’t change their character every few months.
- cognitiveinline 3mo agoThere's a lot people will do for equal quality and faster work at a 100th the cost. Like being OK with character changing and adapting to it. You are not seeing this in your workplace?
- baq 3mo agoIt’s going to even out on both eventually, the diffusion will be painful though.
- nasretdinov 3mo agoI can agree with Dan on two things: LLMs do often produce incorrect results and that it's still useful for productivity when used in moderation. For me the wrong results actually cause some kind of ragebait response so I become much more motivated to learn more about the subject to actually generate correct response. After I've learnt the subject area enough I find I'm better off having LLM review my code instead of writing it. I haven't even begun to try to comprehend how to use fuzzing testing to improve the ability to find bugs, but it sounds really interesting. I've seen mutation testing to be very useful for finding gaps in tests, so I can only imagine that fuzzing + LLMs might produce insane results.
- joshka 3mo ago[flagged]
- jujube3 3mo ago[flagged]
- TokenLens 3mo ago[flagged]
- deleted 3mo ago[deleted]
- jason_s 3mo agoReally interesting article but there's one thing I wanted to point out about the reasons for code review: > If a company were shipping bugs at, say, a hundredth the rate we were at Centaur while relying primarily on review to catch bugs, then I could see their point, but that's not what's happening at the typical software company where people don't want to move away from human review [there might be non quality (as in non bug rate) related reasons to keep human review, such as keeping a high bar for code quality or keeping the codebase human-understandable, which pretty much immediately stops being the case if you let a fleet of agents go wild on a codebase] because of the perceived risk of shipping bugs. I don't agree that the main point of a code review is quality assurance; there are other reasons that have more to do with team convergence and junior staff learning: https://www.embeddedrelated.com/showarticle/807.php https://www.embeddedrelated.com/showarticle/807.php
- deterministic 3mo agoI've read a lot of comments about using AI for coding, and my experience has been very different. I work on large C++ applications used by international airlines. If this software failed, it would make national headlines. Claude Code with Opus 4.8 is great at handling the boring code I don't want to write myself. It gets it right almost every time. However I still review every change and test everything before committing. Trust, but verify.
- amirofy 3mo ago[flagged]