12 ms·
Eight more months of agents
- almostdeadguy 8mo agoIn the past couple days I've become less skeptical of the capabilities of LLMs and now more alarmed by them, contra the author. I think if we as a society continue to accept the development of LLMs and the control of them by the major AI companies there will be massively negative repercussions. And I don't mean repercussions like "a rogue AI will destroy humanity" per se, but these things will potentially cause massive social upheaval, a large amount of negative impacts on mental health and cognition, etc. I think if you see LLMs as powerful but not dangerous you are not being honest.
- gadflyinyoureye 8mo agoI see them as powerful and dangerous. The goal for decades now is to reduce the human population to 500 million. All human technology was pushed to this end, covertly. If we suddenly have a technology that renders white collar workers useless, we will get to that number faster than expected.
- Applejinx 8mo agoI don't believe that is true, but if it WAS true that human technology was covertly pushed to this end: there are people out there who are demanding that this technology come up with social manipulations (using language) to reduce the human population to a SPECIFIC 500 million. Or less. And I don't think it's collar color they're going to be checking against. So I guess I'm saying I agree that this is powerful and dangerous. These are language models, so they're more effective against humans and their languages. And self-preservation, empathy, humanity do not play a role as there is nobody in there to be offended at the notion of intentionally killing more than 9/10 of humanity… for some definitions of humanity, ones I'm sympathetic to.
- NitpickLawyer 8mo agoThere are some good things here: First, we currently have 4 frontier labs, and a bunch of 2nd tier ones following. The fact that we don't have just oAI or just Anthropic or just Google is good in the general sense, I would say. The 4 labs racing each other and trading SotA status for ~a few weeks is good for the end consumer. They keep each other honest and keep the prices down. Imagine if Anthropic could charge 60$ /MTok or oAI could charge 120$ /MTok for their gpt4 style models. They can't in good part because of the competition. Second, there's a bunch of labs / companies that have released and are continuing to release open mdoels. That's as close to "intelligence on tap" as you can get. And those models are ~6-12 months behind the SotA models, depending on your usecase. Even though the labs have largely different incentives to do so, a lot of them are still releasing open models. Hopefully that continues to hold. So not all control will be in the hands of big tech, even if the "best" will still be theirs. At some point "good enough" is fine. There's also the thing about geopolitics being involved in this. So far we've seen the EU jumping the gun on regulation, and we're kinda sorta paying for it. Everyone is still confused about what can or cannot be done in the EU. The US seems to be waiting to see what happens, and China will do whatever they do. The worst thing that can happen is that at some point the big players (Anthropic is the main driver) push for regulatory capture. That would really suck. Thankfully atm there's this lingering thinking that "if we do it, the others won't so we'll be on the back foot". Hopefully this holds, at least until the "good enough" from above is out :)
- almostdeadguy 8mo agoI'm not just concerned about control by one company, I'm concerned by control for the profit motive, and probably concerned about the wisdom of using these things for anything except extremely limited use cases (breakthrough scientific research, etc.). I think tech people have a bad tendency of viewing this through the lens of platform wars type stakes, and there are much bigger problems with AI. The fact that an alarming number of ex-and-current Anthropic people I've met think the world is going to end is something we should take heed of! The AI labs started down this path using the Manhattan Project as a metaphor and guess what? It's a good metaphor and we should embrace most of the wider implications of that (though I'd love to avoid all the MAD/cold war bullshit this time).
- hasperdi 8mo ago> It sounds like someone saying power tools should be outlawed in carpentry. I see this a lot here
- sp33der89 8mo agoAll metaphors break down at a certain point, but power tools and generative AI/LLMs being compared feels like somebody is romanticizing the art of programming a bit too much. Copyright law, education, just the sheer scale of things changing because of LLMs are some things off the top of my head why "power tools vs carpentry" is a bad analogy.
- Keyframe 8mo agoif that someone is clumsy, had an active war going on against basic tools before, and wandered into the carpentry from completely different area, then power tools might be a bad idea.
- deleted 8mo ago[deleted]
- wasmainiac 8mo agoYes because A tech-bro AIs dream is hundreds of thousands of developers being let go and replacing them with no code tools. Sure, replace me with AI, but I better get royalties on my public contributions. I like many other developers have kids and other responsibilities to pay for. We did not share our work publicly to be replaced. The same way I did not lend my neighbour my car so he could run me over, that was implicit.
- senordevnyc 8mo agoWe’ve been doing this to other professions for half a century. Live by the sword, die by the sword.
- wasmainiac 8mo agoAre you really going to appeal to nature here? As a species we are basically magic. We have gone to the moon, cured countless diseases and many other amazing things, but not to stop being taken advantage of? Is it too high of an ask?
- happytoexplain 8mo agoI don't trust the idea of "not getting", "not understanding", or "being out of touch" with anti-LLM (or pro-LLM) sentiment. There is nothing complicated about this divide. The pros and cons are both as plain as anything has ever been. You can disagree - even strongly - with either side. You can't "not understand".
- slfnflctd 8mo ago> There is nothing complicated about this divide [...] You can't "not understand" I beg to differ. There are a whole lot of folks with astonishingly incomplete understanding about all the facts here who are going to continue to make things very, very complicated. Disagreement is meaningless when the relevant parties are not working from the same assumption of basic knowledge.
- illusive4080 8mo agoBolstering your point, check out the comments in this thread: https://www.reddit.com/r/rust/comments/1qy9dcs/who_has_completely_sworn_off_including_llm/ https://www.reddit.com/r/rust/comments/1qy9dcs/who_has_compl... There’s a lot of unwillingness to even attempt to try the tools.
- cruffle_duffle 8mo agoThose people are absolutely going to get left in the dust. In the hands of a skilled dev, these things are massive force multipliers.
- dang 8mo agoRelated. Others? How I program with agents - https://news.ycombinator.com/item?id=44221655 https://news.ycombinator.com/item?id=44221655 - June 2025 (295 comments)
- dirkc 8mo ago> Using anything other than the frontier models is actively harmful If that is true, why should one invest in learning now rather than waiting for 8 months to learn whatever is the frontier model then?
- fusslo 8mo agosnarky answer: so you can be that 'AI guy' at your office that everyone avoids in the snackroom
- ej88 8mo agoIt's not like you need to take a course. The frontier models are the best, just using them and their harnesses and figuring out what works for your use case is the 'investing in learning'.
- recursive 8mo agoHow could it be actively harmful if it wasn't harmful last month when it was the frontier model?
- deleted 8mo ago[deleted]
- senko 8mo agoBecause you might want to use LLMs now. If not, it's definitely better to not chase the hype - ignore the whole shebang. But if you do want to use LLMs for coding now, not using the best models just doesn't make sense.
- jonas21 8mo agoSo that you can be using the current frontier model for the next 8 months instead of twiddling your thumbs waiting for the next one to come out? I think you (and others) might be misunderstanding his statement a bit. He's not saying that using an old model is harmful in the sense that it outputs bad code -- he's saying it's harmful because some of the lessons you learn will be out of date and not apply to the latest models. So yes, if you use current frontier models, you'll need to recalibrate and unlearn a few things when the next generation comes out. But in the meantime, you will have gotten 8 months (or however long it takes) of value out of the current generation.
- dmk 8mo agoThe real insight buried in here is "build what programmers love and everyone will follow." If every user has an agent that can write code against your product, your API docs become your actual product. That's a massive shift.
- anthuswilliams 8mo agoI'm very much looking forward to this shift. It is SO MUCH more pro-consumer than the existing SaaS model. Right now every app feels like a walled garden, with broken UX, constant redesigns, enormous amounts of telemetry and user manipulation. It feels like every time I ask for programmatic access to SaaS tools in order to simplify a workflow, I get stuck in endless meetings with product managers trying to "understand my use case", even for products explicitly marketed to programmers. Using agents that interact with APIs represents people being able to own their user experience more. Why not craft a frontend that behaves exactly the the way YOU want it to, tailor made for YOUR work, abstracting the set of products you are using and focusing only on the actual relevant bits of the work you are doing? Maybe a downside might be that there is more explicit metering of use in these products instead of the per-user licensing that is common today. But the upside is there is so much less scope for engagement-hacking, dark patterns, useless upselling, and so on.
- pjc50 8mo ago> Right now every app feels like a walled garden, with broken UX, constant redesigns, enormous amounts of telemetry and user manipulation OK, but: that's an economic situation. > so much less scope for engagement-hacking, dark patterns, useless upselling, and so on. Right, so there's less profit in it. To me it seems this will make the market more adversarial, not less. Increasing amounts of effort will be expended to prevent LLMs interacting with your software or web pages. Or in some cases exploit the user's agentic LLM to make a bad decision on their behalf.
- 13pixels 8mo agothe "exploit the user's agentic LLM" angle is underappreciated imo. we already see prompt injection attacks in the wild -- hidden text on web pages that tells the agent to do things the user didn't ask for. now scale that to every e-commerce site, every SaaS onboarding flow, every comparison page. it's basically SEO all over again but worse, because the attack surface is the user's own decision-making proxy. at least with google you could see the search results and decide yourself. when your agent just picks a vendor for you based on what it "found," the incentive to manipulate that process is enormous. we're going to need something like a trust layer between agents and the services they interact with. otherwise it's just an arms race between agent-facing dark patterns and whatever defenses the model providers build in.
- dagss 8mo agoBut if you try some penny-saving cheap model like Sonnet [..bad things..]. [Better] pay through the nose for Opus. After blowing $800 of my bootstrap startup funds for Cursor with Opus for myself in a very productive January I figured I had to try to change things up... so this month I'm jumping between Claude Code and Cursor, sometimes writing the plans and having the conversation in Cursor and dump the implementation plan into Claude. Opus in Cursor is just so much more responsive and easy to talk to, compared to Opus in Claude. Cursor has this "Auto" mode which feels like it has very liberal limits (amortized cost I guess) that I'm also trying to use more, but -- I don't really like to flip a coin and if it lands up head then waste half hour discovering the LLM made a mess the LLM and try again forcing the model. Perhaps in March I'll bite the bullet and take this authors advice.
- written-beyond 8mo agoJust use Codex 5.3 in codex cli, the $20/mo plan is basically limitless at least for me and I keep reasoning efforts high. You can enjoy it while it lasts, OpenAI is being very liberal with their limits because of CC eating their lunch rn.
- rmonvfer 8mo agoYeah, I can’t recommend gpt-5.3-codex enough, it’s great! I’ve been using it with the new macOS app and I’m impressed. I’ve always been a Claude Code guy and I find myself using codex more and more. Opus is still much nicer explaining issues and walking me through implementations but codex is faster (even with xhigh effort) and gets the job done 95% of the time. I was spending unholy amounts of money and tokens (subsidized cloud credits tho) forcing Opus for everything but I’m very happy with this new setup. I’ve also experimented with OpenCode and their Zen subscription to test Kimi K2.5 an similar models and they also seem like a very good alternative for some tasks. What I cannot stand tho is using sonnet directly (it’s fine as a subagent), I’ve found it to be hard to control and doesn’t follow detailed instructions.
- jorl17 8mo agoOut of curiosity, what’s your flow? Do you have codex write plans to markdown files? Just chat? What languages or frameworks do you use? I’m an avid cursor user (with opus), and have been trying alternatives recently. Codex has been an immense letdown. I think I was too spoiled by cursor’s UX and internal planning prompt. It’s incredibly slow, produces terribly verbose and over-complicated code (unless I use high or xhigh, which are even slower), and missed a lot of details. Python/django and react frontend. For the first time I felt like I could relate to those people who say it doesn’t make them faster,” because they have to keep fixing the agent’s shot, never felt that with opus 4.5 and 4.6 and cursor
- emmawirt 8mo agoCurious what you mean by "agent harness" here... are you distinguishing between true autonomous agents (model decides next step) vs workflows that use LLMs at specific nodes? I've found the latter dramatically more reliable for anything beyond prototyping, which makes me wonder if the "model improvement" is partly better prompting and scaffolding.
- crawshaw 8mo agoHi, author here. I mean the piece of code that calls the model and executes the tool calls. My colleague Philip calls it “9 lines of code”: https://sketch.dev/blog/agent-loop https://sketch.dev/blog/agent-loop We have built two of them now, and clearly the state of the art here can be improved. But it is hard to push too much on this while the models keep improving.
- tiny-automates 8mo agothe harness being "9 lines of code" is deceptive in the same way a web server is "just accept connections and serve files." the hard part isn't the loop itself — it's everything around failure recovery. when a browser agent misclicks, loads a page that renders differently than expected, or hits a CAPTCHA mid-flow, the 9-line loop just retries blindly. the real harness innovation is going to be in structured state checkpointing so the agent can backtrack to the last known-good state instead of restarting the whole task. that's where the gap between "works in a demo" and "works on the 50th run" lives.
- rahimnathwani 8mo agoAn agent harness is what enables the user to seamlessly interact with both a model and tool calls. Claude Code is an agent harness. ┌────────────────────────────┐ │ User │ └──────────────┬─────────────┘ │ ▼ ┌────────────────────────────┐ │ Agent Harness │ │ (software interface) │ └──────┬──────────────┬──────┘ │ │ ▼ ▼ ┌────────────┐ ┌────────────┐ │ Models │ │ Tools │ └────────────┘ └────────────┘ Here's an example of a harness with less code: https://github.com/badlogic/pi-mono/blob/fdcd9ab783104285764a4b770c39b7811be5f570/packages/agent/src/agent-loop.ts#L104 https://github.com/badlogic/pi-mono/blob/fdcd9ab783104285764...
- Herring 8mo ago> In 2000, less than one percent lived on farms and 1% of workers are in agriculture. That was a net benefit to the world, that we all don't have to work to eat. The jury's still out on that one, because climate change is an existential risk.
- gradus_ad 8mo agoExistential? Maybe to beachfront property owners
- Herring 8mo agoEverybody gangsta until the permafrost starts leaking massive amounts of methane.
- esafak 8mo agoDid you know that 10% of the world's population lives in coastal zones at low elevations?
- ares623 8mo agoBut they'll just disappear into thin air peacefully when that happens right? It's not like they're gonna fight tooth and nail to find a place to survive, that'd be rude.
- xyzsparetimexyz 8mo agoIdiot
- joefourier 8mo agoThe author is correct in that agents are becoming more and more capable and that you don't need the IDE to the same extent, but I don't see that as good. I find that IDE-based agentic programming actually encourages you to read and understand your codebase as opposed to CLI-based workflows. It's so much easier to flip through files, review the changes it made, or highlight a specific function and give it to the agent, as opposed to through the CLI where you usually just give it an entire file by typing the name, and often you just pray that it manages to find the context by itself. My prompts in Cursor are generally a lot more specific and I get more surgical results than with Claude Code in the terminal purely because of the convenience of the UX. But secondly, there's an entire field of LLM-assisted coding that's being almost entirely neglected and that's code autocomplete models. Fundamentally they're the same technology as agents and should be doing the same thing: indexing your code in the background, filtering the context, etc, but there's much less attention and it does feel like the models are stagnating. I find that very unfortunate. Compare the two workflows: With a normal coding agent, you write your prompt, then you have to at least a full minute for the result (generally more, depending on the task), breaking your flow and forcing you to task-switch. Then it gives you a giant mass of code and of course 99% of the time you just approve and test it because it's a slog to read through what it did. If it doesn't work as intended, you get angry at the model, retry your prompt, spending a larger amount of tokens the longer your chat history. But with LLM-powered auto-complete, when you want, say, a function to do X, you write your comment describing it first, just like you should if you were writing it yourself. You instantly see a small section of code and if it's not what you want, you can alter your comment. Even if it's not 100% correct, multi-line autocomplete is great because you approve it line by line and can stop when it gets to the incorrect parts, and you're not forced to task switch and you don't lose your concentration, that great sense of "flow". Fundamentally it's not that different from agentic coding - except instead of prompting in a chatbox, you write comments in the files directly. But I much prefer the quick feedback loop, the ability to ignore outputs you don't want, and the fact that I don't feel like I'm losing track of what my code is doing.
- coffeefirst 8mo agoThe other thing about non-agent workflows is they’re much, much less compute intensive. This is going to matter.
- dsign 8mo agoLook, I'm very negative about this AI thing. I think there is a great chance it will lead to something terrible and we will all die, or worse. But on the other hand, we are all going to die anyway. Some of us, the lucky ones, will die of a heart attack and will learn of our imminent demise in the second it happens, or not at all. The rest of us will have it worse. It has always been like that, and it has only gotten more devastating since we started wearing clothes and stopped being eaten alive by a savanna crocodile or freezing to death during the first snowfall of winter. But if AI keeps getting better at code, it will produce entire in-silico simulation workflows to test new drugs or even to design synthetic life (which, again, could make us all die, or worse). Yet there is a tiny, tiny chance we will use it to fix some of the darkest aspects of human existence. I will take that.
- xyzsparetimexyz 8mo agoThat's stupid. If you genuinely think that there's a great chance AI will kill us all, you wouldn't spin the wheel just for some small vague chance that it doesn't and something good (what exactly, nobody knows) will happen
- habinero 8mo agoA lot of very silly people have convinced themselves and you of this, but it is not true, was never true, and is never going to be true. We have a lot of actual problems to deal with that aren't telling ghost stories about sand. Focus on those.
- Krei-se 8mo agoThe author has a github.
- 63 8mo agoI have no problem with experienced senior devs using agents to write good code faster. What I have a problem with is inexperienced "vibecoders" who don't care to learn and instead use agents to write awful buggy code that will make the product harder to build on even for the agents. It used to be that lack of a basic understanding of the system was a barrier for people, but now it's not, so we're flooded with code written by imperfect models conducted by people who don't know good from bad.
- codebolt 8mo agoWhere are you encountering all this slop code? At my work we use LLMs heavily and I don't see this issue. Maybe I'm just lucky that my colleagues all have Uni degrees in CS and at least a few years experience.
- post-it 8mo ago> Maybe I'm just lucky that my colleagues all have Uni degrees in CS and at least a few years experience. That's why. I was using Claude the other day to greenfield a side project and it wanted to do some important logic on the frontend that would have allowed unauthenticated users to write into my database. It was easy to spot for me, because I've been writing software for years, and it only took a single prompt to fix. But a vibe coder wouldn't have caught it and hackers would've pwned their webapp.
- giancarlostoro 8mo agoYou can also ask Claude to review all the code for security issues and code smells, you'd be surprised what it finds. We all write insecure code in our first pass through if we're too focused on getting the proof of concept worked out, security isnt always the very 1st thing coded, maybe its the very next thing, maybe it comes 10 changes later.
- blibble 8mo ago> We all write insecure code in our first pass through no, we don't
- post-it 8mo ago> Agent harnesses have not improved much since then. There are things Sketch could do well six months ago that the most popular agents cannot do today. I think this is a neglected area that will see a lot of development in the near future. I think that even if development on AI models stopped today - if no new model was ever trained again - there are still decades of innovation ahead of us in harnessing the models we already have. Consider ChatGPT: the first release relied entirely on its training data to answer questions. Today, it typically does a few Google searches and summarizes the results. The model has improved, but so has the way we use it.
- kevmo314 8mo agoReally? I hardly think it's neglected. The Claude Code harness is the only reason I come back to it. I've tried Claude via OpenCode or others and it doesn't work as well for me. If anything, I would even argue that prior to 4.6, the main reason Opus 4.5 felt like it improved over months was the harness.
- tiny-automates 8mo agoagreed, and i'd go further - the harness is where evaluation actually happens, not in some separate benchmark suite. rhe model doesn't know if it succeeded at a web task. the harness has to verify DOM state, check that the right element was clicked, confirm the page transitioned correctly. right now most harnesses just check "did the model say it was done" which is why pass rates on benchmarks don't translate to production reliability. the interesting harness work is building verification into the loop itself, not as an afterthought.
- uludag 8mo ago> I am having more fun programming than I ever have, because so many more of the programs I wish I could find the time to write actually exist. I wish I could share this joy with the people who are fearful about the changes agents are bringing. It might be just me but this reads as very tone deaf. From my perspective, CEOs are seething at the mouth to make as many developers redundant as possible, not being shy about this desire. (I don't see this at all as inevitable, but tech leaders have made their position clear) Like, imagine the smugness of some 18th century "CEO" telling an artisan, despite the fact that he'l be resigned to working in horrific conditions at a factory, to not worry and think of all the mass produced consumer goods he may enjoy one day. It's not at all a stretch of the imagination that current tech workers may be in a very precarious situation. All the slopware in the world wouldn't console them.
- overgard 8mo agoI bought Steve Yegge's "Vibe Coding" book. I think I'm about 1/4th of the way through it or so. One thing that surprised me is there's this naivete on display that workers are going to be the ones to reap the benefits of this. Like, Steve was using an example of being able to direct the agent while doing leisure activities (never mind that Steve is more of an executive/thought leader in this company, and, prior to LLMs, seemed to be out of the business of writing code). That's a nice snapshot of a reality that isn't going to persist.. While the idea of programmers working two hours a day and spending the rest of it with their family seems sunny, that's absolutely not how business is going to treat it. Thought experiment... CEO has a team of 8 engineers. They do some experiments with AI, and they discover that their engineers are 2x more effective on average . What does the CEO do? a) Change the workweek to 4 hours a day so that all the engineers have better work/life balance since the same amount of work is being done. b) Fire half the engineers, make the 4 remaining guys pick up the slack, rinse and repeat until there's one guy left? Like, come on. There's pushback on this stuff not because the technology is bad, (although it's overhyped), but because the no sane person trusts our current economic system to provide anything resembling humane treatment of workers. The super rich are perfectly fine seeing half the population become unemployed, as far as I can tell, as long as their stock numbers go up.
- 8mo ago
- monus 8mo ago> Along the way I have developed a programming philosophy I now apply to everything: the best software for an agent is whatever is best for a programmer. Not a plug but really that’s exactly why we’re building sandboxes for agents with local laptop quality. Starting with remote xcode+sim sandboxes for iOS, high mem sandbox with Android Emulator on GPU accel for Android. No machine allocation but composable sandboxes that make up a developer persona’s laptop. If interested, a quick demo here https://www.loom.com/share/c0c618ed756d46d39f0e20c7feec996d https://www.loom.com/share/c0c618ed756d46d39f0e20c7feec996d muvaf[at]limrun[dot]com
- panny 8mo agoI see a lot of people here saying things like: >ah, they're so dumb, they don't get it, the anti-LLM people This is one of the reasons I see AI failing in the short term. If I call you an idiot, are you more or less likely to be open minded and try what I'm selling? AI isn't making money, 95% of companies are failing with AI https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/ https://fortune.com/2025/08/18/mit-report-95-percent-generat... I mean, your AIs might be a lot more powerful if it was generating money, but that's not happening. I guess being condescending to the 95% of potential buyers isn't really working out.
- jeffrallen 8mo agoListen to this guy. I've been using his code for a long time, and it works. I am a happy customer of his service, and it works. I listen to his advice and it works.
- gip 8mo ago> In 2026, I don't use an IDE any more. I don't think it is the best way to look at it. I think that now every team has the power to build and maintain an internal agent (tool + UX) to manager software products. I don't necessarily think that chat-only is enough except for small projects, so teams will build agent that gives them access to the level of abstraction that works best. It's a data point but this weekend (e.g. in 2 days) I build a desktop + web agent that is able to help me reason on system design and code. Built with Codex powered by the Codex SDK. It is high quality. I've been a software engineer and director of engineering for 10 years. I'm blown away.
- sarchertech 8mo agoI’m not saying this is definitely a bot. However, this is the 7th time I’ve read a post and thought it might be an OpenAI promotion bot, clicked on the username, and noticed that the account was created in 2011. I have yet to do this and see any other year. Was there someone who bought a ton of accounts in 2011 to farm them out? A data breach? Was 2011 just a very big year for new users? (My own account is from 2011)
- loveparade 8mo agoIt's definitely a bot, just like probably around 10% of comments on HN at this point, and the majority of upvotes. And it's only increasing. Calling it bot is a bit dismissive though. It's an agent!
- redanddead 8mo agoit is giving a very agentic vibe
- gip 8mo agoCare to have a phone call with who you call a bot tonight? If so, send a DM on twitter to @edfixyz with your phone number and I will call you immediately. Or give me your twitter handle. I'm tired of that BS - when people don't like what you write they call you a bot.
- 8mo ago
- dmos62 8mo ago> the best software for an agent is whatever is best for a programmer My conclusion as well. It feels paradoxical, maybe because on some level I still think of an LLM as some weird gadget, not a coworker. Context ephemerality is more or less the only veritable difference from a human programmer, I'd say. And, even then, context introduction with LLMs is a speedrun of how you'd do it with new human members of a project. Awesome times we live in.
- hoistbypetard 8mo ago> To me that statement is as obvious as "water is wet". Water is not wet. Water makes things wet. Perhaps the inaccuracy of that statement should be taken as a hint that the other statements that you hold on the same level are worthy of reconsideration.
- baq 8mo agoThe good old classic technically correct and completely besides the point observation.
- pseudosavant 8mo agoYes, technically HN is full of these kinds of corrections... but HN isn't actually wet.
- swyx 8mo agousername checks out
- xyzsparetimexyz 8mo ago> Pay through the nose for Opus or GPT-7.9-xhigh-with-cheese. Don't worry, it's only for a few years. > You have to turn off the sandbox, which means you have to provide your own sandbox. I have tried just about everything and I highly recommend: use a fresh VM. > I am extremely out of touch with anti-LLM arguments 'Just pay out the arse and run models without a sandbox or in some annoying VM just to see them fail. Wait, some people are against this?'
- dude250711 8mo agoWell, not 'against' per se, just watching LLM-enthusiasts tumble in the mud for now. Though I have heard that if I don't jump into the mud this instance, I will be left behind apparently for some reason. So you either get left behind or get a muddy behind, your choice.
- pferde 8mo agoEverybody keeps saying the models are getting better, the tooling is getting better, people are discovering better practices... So why not just wait out this insane initial phase, and if anything is left standing afterwards and proves itself, just learn that.
- kjksf 8mo agoBecause there's nothing to learn. "learning" to use claude code is less effort than learning how to use the basics of git. They provide value today so I'm using them today.
- Expurple 8mo agoThis. If you use a modern frontier model like Opus 4.5, there's nothing to learn. No special prompting techniques. You give it a task, and most of the time it's capable of solving a big chunk quickly. You still need to babysit it, review its plan/code and make adjustments. But that's already faster than achieving the same results manually. Especially when you're at low energy levels and can't bring yourself to look into a task and figure it out from zero.
- imron 8mo ago> By far the greatest IDE I have ever used was Visual Studio C++ 6.0 on Windows 2000 Visual C++ 6 was incredible! My favourite IDE of all time too.
- EdwardDiego 8mo agoPoor fellow has never used IntelliJ IDEA.
- imron 8mo agoYes I have.
- 0xbadcafebee 8mo agoLocal models are decent now. Qwen3 coder is pretty good and decent speed. I use smaller models (qwen2.5:1.5b) with keyboard shortcuts and speech to text to ask for man page entries, and get 'em back faster than my internet connection and a "robust" frontier model does. And web search/RAG hides a multitude of sins. "Using anything other than the frontier models is actively harmful" - so how come I'm getting solid results from Copilot and Haiku/Flash? Observe, Orient, Decide, Act, Review, Modify, Repeat. Loops with fancy heuristics, optimized prompts, and decent tools, have good results with most models released in the past year.
- mattmanser 8mo agoHave you used the frontier models recently? It's hard to communicate the difference the last 6 months has seen. We're at the point where copilot is irrelevant. Your way of working is irrelevant. Because that's not how you interact with coding AIs anymore, you're chatting with them about the code outside the IDE.
- timr 8mo ago> Have you used the frontier models recently? Yes. > It's hard to communicate the difference the last 6 months has seen. No, it isn't. The hypebeast discovered Claude code, but hasn't yet realized that the "let the model burn tokens with access to a shell" part is the key innovation, not the model itself. I can (and do) use GH Copilot's "agent" mode with older generation models, and it's fine. There's no step function of improvement from one model to another, though there are always specific situations where one outperforms. My current go-to model for "sit and spin" mode is actually Grok, and I will splurge for tokens when that doesn't work. Tools and skills and blahblahblah are nice to have (and in fact, part of GH Copilot now), but not at all core to the process.
- gjulianm 8mo agoHonestly, I've been using the frontier models and I'm not sure where people are seeing these massive improvements. It's not that they're bad, it's just that I don't see that much of an improvement the last 6 months. They're so inconsistent that it's hard to have a clear idea of what's happening. I usually switch between models and I don't see either those massive differences either. Not to mention that sometimes models regress in certain aspects (e.g., I've seen later models that tend to "think" more and end up at the same result but taking far more time and tokens).
- symfrog 8mo agoAny sufficiently complicated LLM generated program contains an ad hoc, informally-specified, bug-ridden, slow implementation of half of an open source project.
- kakacik 8mo agoWe had an effort recently where one much more experienced dev from our company ran Claude on our oldish codebase for one system, with the goal of transforming it into newer structure, newer libraries etc. while preserving various built in functionalities. Not the first time this guy did such a thing and he is supposed to be an expert. I took a look at the result and its maybe half of stuff missing completely, rest is cryptic. I know that codebase by heart since I created it. From my 20+ years of experience correcting all this would take way more effort than manual rewrite from scratch by a senior. Suffice to say thats not what upper management wants to hear, llm adoption often became one of their yearly targets to be evaluated against. So we have a hammer and looking for nails to bend and crook. Suffice to say this effort led nowhere since we have other high priority goals, for now. Smaller things here & there, why not. Bigger efforts, so far sawed-off 2-barrel shotgun loaded with buckshot right into both feet.
- kjksf 8mo agoNot to take away from your experience but to offer a counterpoint. I used claude code to port rust pdb parsing library to typescript. My SumatraPDF is a large C++ app and I wanted visibility into where does the size of functions / data go, layout of classes. So I wanted to build a tool to dump info out of a PDB. But I have been diagnosed with extreme case of Rustophobiatis so I just can't touch rust code. Hence, the port to typescript. With my assistance it did the work in an afternoon and did it well. The code worked. I ran it against large PDB from SumatraPDF and it matched the output of other tools. In a way porting from one language to another is extreme case of refactoring and Claude did it very well. I think that in general (your experience notwithstanding) Claude Caude is excellent at refactorings. Here are 3 refactorings from SumatraPDF where I asked claude code to simplify code written by a human: https://github.com/sumatrapdfreader/sumatrapdf/commit/a472d32048da750b45a9b39ba5da3e649d6165c5 https://github.com/sumatrapdfreader/sumatrapdf/commit/a472d3... https://github.com/sumatrapdfreader/sumatrapdf/commit/5624aa73729b0a5aaea8bc4c5d95328d5f7ec6ff https://github.com/sumatrapdfreader/sumatrapdf/commit/5624aa... https://github.com/sumatrapdfreader/sumatrapdf/commit/a40bc91e4da2338bf8cc8748277416db058ee58d https://github.com/sumatrapdfreader/sumatrapdf/commit/a40bc9... I hope you agree the code written by Claude is better than the code written by a human. Granted, those are small changes but I think it generalizes into bigger changes. I have few refactorings in mind I wanted to do for a long time and maybe with Claude they will finally be feasible (they were not feasible before only because I don't have infinite amount of time to do everything I want to do).
- entropyneur 8mo ago> I deeply appreciate hand-tool carpentry and mastery of the art, but people need houses and framing teams should obviously have skillsaws. Where are all the new houses? I admit I am not a bleeding edge seeker when it comes to software consumption, but surely a 10x increase in the industry output would be noticeable to anyone?
- xyzzy123 8mo agoOrg processes have not changed. Lots of the devs I know are enjoying the speedup on mundane work, consuming it as a temporary lifestyle surplus until everything else catches up. You can't saw faster than the wood arrives. Also the layout of the whole job site is now wrong and the council approvals were the actual bottleneck to how many houses could be built in the first place... :/
- gbuk2013 8mo agoBasically this. My last several tickets were HITL coding with AI for several hours and then waiting 1-2 days while the code worked its way through PR and CI/CD process. Coding speed was never really a bottleneck anywhere I have worked - it’s all the processes around it that take the most time and AI doesn’t help that much there.
- dent9 8mo agoTrue story; I wanted to make a tiny update to our CI / CD to upload copies of some artifacts to S3. It took 1min for the LLM to remind me of the correct syntax in aws cli to do the upload and the syntax to plug it into our GitHub Actions. It then took me the next 3 hours to figure out which IAMs needed to be updated in order to allow the upload before it was revealed that Actually uploading to the S3 requires the company IT to adjust bucket policies and this requires filing a ticket with IT and waiting 1-5 business days for a response then potentially having a call with them to discuss the change and justify why we need it. So now it's four days later and I still can't push to S3. AI reduced this from a 5-day process to a 4.9-day process
- bunderbunder 8mo ago
- only2people 8mo ago>this is why I'm building My clipart folder of that kid with the lolipop continues to stay relevant
- ares623 8mo agoguys it's an ad
- indigodaddy 8mo agoNah, it's not, really
- nickcw 8mo ago> Some believe AI Super-intelligence is just around the corner (for good or evil). Others believe we're mistaking philosophical zombies for true intelligence, and speedrunning our own brainrot Not sure which camp I'm in, but I enjoyed the imagery.
- MrSandingMan 8mo ago> In 2000, less than one percent lived on farms and 1% of workers are in agriculture. That was a net benefit to the world, that we all don't have to work to eat. Not obvious > To me that statement is as obvious as "water is wet". Well... is water *wet* or does it *wet things*? So not obvious either. I'm really dubious when reading posts posing some things as obvious or trivial. In general they are not.
- jFriedensreich 8mo agoIts funny how many variations of meaning people assign to agent related terms. Conflating agent with cli and as opposite spectrum of ide is a new one i did not encounter before. I run agents with vscode-server also in a vm and would not give up the ability to have a proper gui anytime i feel like and also being able to switch seamless between more autonomous operation and more interactive seems useful at any level.
- piokoch 8mo ago"In 2026, I don't use an IDE any more." Just a question? What IDE feature is obsolete now? Ability to navigate the code? Integration with database, Docker, JIRA, Github (like having PR comments available, listed, etc), Git? Working with remote files? Building the project? Yes, I can ask copilot to build my project and verify tests results, but it will eat a lot of tokens and added value is almost none.
- Expurple 8mo ago> I can ask copilot to build my project and verify tests results, but [..] added value is almost none. The added value is that it can iterate autonomously and finish tasks that it can't one-shot in its first code edit. Which is basically all tasks that I assign to Copilot. The added value is that I get to review fully-baked PRs that meet some bar of quality. Just like I don't review human PRs if they don't pass CI. Fully agree on IDEs, though. I absolutely still need an IDE to iterate on PRs, review them, and tweak them manually. I find VSCode+Copilot to be very good for this workflow. I'm not into vibe coding.
- Havoc 8mo agoDisagree with the point about anything less than opus being harmful to learning. Much of my learning still requires experimentation - including lots of token volume so hitting limits is a problem. And secondly I’m looking for workflows that build the thing without needing to be at the absolute edge of the LLM capability. Thats where fragility and unpredictability live. Where a new model with slightly different personality is released and it breaks everything. I’d rather have flow that is simple and idiot proof that doesn’t fall apart at the first sign of non-bleeding edge tokens. That means skipping the gains from something opus could one shot ofc but that’s acceptable to me
- thegrim000 8mo ago"Eight more months of Bitcoin. It's usage continues to dramatically expand. The amount of transactions is increasing exponentially. Soon, fiat currencies will collapse, all replaced by Bitcoin transactions. If you haven't converted your assets over to Bitcoin you're going to be left behind and lose it all. I can't even understand people that don't see the obvious technical superiority of Bitcoin, such people are going to go through rough times."
- webdevver 8mo ago>I wish I could share this joy with the people who are fearful about the changes agents are bringing. The 'fear' is about losing ones livelihood and getting locked out of homeownership and financial security. its not complicated. life is actually largely determined by your access to capital, despite whatever fresh coping strategy the afflicted (and the afflicting) like to peddle. the quality of life versus capital availability is very non-linear. there is a step-change around the $500k mark where you reach 'orbital velocity', where as long as you dont suffer severe misfortune or make mistakes, you will start accelerating upwards (albeit very slowly.) under that line, you are constantly having to fight 'gravity'. basically everyone in tech is openly or quietly aiming to get there, and LLMs have made that trek ever more precarious than before.
- simondoubleu 8mo ago[dead]
- mtlynch 8mo ago> Along the way I have developed a programming philosophy I now apply to everything: the best software for an agent is whatever is best for a programmer. I agree with this and I think it's funny to see people publish best practices for working with AI that are like, "Write a clear spec. Have a style guide. Use automated tests." I'm not convinced it's 100% true because I think there are code patterns that AI handles better than humans and vice versa. But I think it's true enough to use as a guiding philosophy.
- vedhant 8mo agoThe sandboxing pain is real. Sadly, a new VM seems like the most simple and viable solution. I don't think the masses are doing any sandboxing at all. We really need a sandbox solution that is sort of dynamic and doesn't pester the user with allow/deny requests. It has to be intelligent and keep up with the llm agents.
- conartist6 8mo agoIDEs are going to come roaring back. As the author says, there's nothing wrong with the idea of the IDE. Of course you want to be using the best, most powerful tools! AI showed us that our current-gen text-editor-first IDEs are massively underserving the needs of the public, yes, but it didn't really solve that problem. We still need better IDEs! What has changed is that we now understand how badly we need them. (source: I am an IDE author)
- sandgrownun 8mo agoI agree with his assessment up until this point in time, it is where we currently are. But it seems to me there is still a large chunk of engineers who don't extrapolate capability out to the engineer being taken out of the loop completely. Imo, it happens in fairly short order. 2-3 years.
- symfrog 8mo agoOn what basis are you making that prediction?
- jmull 8mo agoRegarding the shift away from time spent on agriculture over the last century or so.. > That was a net benefit to the world, that we all don't have to work to eat. I’m pretty sure most all of us are still working to have food to eat and shelter for ourselves and our families. Also, while the on-going industrial and technological revolution has certainly brought benefits, it’s an open question as to whether it will turn out to be a net benefit. There’s a large-scale tragedy of the commons experiment playing out and it’s hard to say what the result will be.
- redkoala 8mo agoWhat are those 3 sentences that the author typed to replicate Stripe for his situation?
- gurjeet 8mo ago> By far the greatest IDE I have ever used was Visual Studio C++ 6.0 on Windows 2000. I have never felt like a toolchain was so complete and consistent with its environment as there. +1. I've tried many times, and failed, to replicate the joy of using that toolchain.
- dent9 8mo ago> I am extremely out of touch with anti-LLM arguments Wow I know that feel. I'm here using LLM for daily work and even hobbies in very conservative manners and didn't think much of it. Now when I have casual discussions with other folks, especially non-tech people, the visceral hatred I get for even mentioning AI and the fact that I use it is insane. There's like an entire sub group of people who are so out of touch with these tools they think they're the devil like the anti-GMO crazies and the PETA psychos.
- zerotolerance 8mo agoThere is a lot of "my" floating around in this article. I always love getting peeks into experiences with this sort of thing, but I think the "mys" highlight something I've seen every day. These agents are really great at bespoke personal flows that build up a TON of almost personal tribal knowledge about how things get done if there is any consistency to those flows at all. Doing this in larger theaters is much more difficult because tribal knowledge is death for larger teams. It drives up the cost of everything which is why individuals or extremely new small teams feel so much more productive. Everything is new here and consistency doesn't matter yet.