26 ms·
A few random notes from Claude coding quite a bit last few weeks
https://xcancel.com/karpathy/status/2015883857489522876 https://xcancel.com/karpathy/status/2015883857489522876
- wkh129857 8mo ago[flagged]
- soganess 8mo ago"addict" Great idea! Le's pathalogize another thing! I love quickly othering whole concepts and putting them in my brain's "bad" box so I can feel superior.
- reducesuffering 8mo agohttps://github.com/karpathy/nanochat https://github.com/karpathy/nanochat https://github.com/karpathy/llm.c https://github.com/karpathy/llm.c The proof is in the pudding. Let's see your code
- jackling 8mo agoI don't agree with the parent commenters characterization of Karpathy, but these projects are just simple toy projects. They're educational material, not production level software.
- rvz 8mo agoYou just proved the parent’s point. He said “…who has never written any production software…” yet you show toy projects instead. Well done.
- lomase 8mo ago[dead]
- yojat661 8mo agoI don't know if it's fair to call him an ai addict or deduce that his ego is bruised. But I do wonder whether karpathy's agentic llm experiences are based on actual production code or pet projects. Based on a few videos I have seen of his, I am guessing it's the latter. Also, he is a research scientist (probably a great one), not a software developer. I agree with the op that karpathy should not be given much attention in this topic i.e llms for software development.
- nadis 8mo agoThe section on IDEs/agent swarms/fallibility resonated a lot for me; I haven't gone quite as far as Karpathy in terms of power usage of Claude Code, but some of the shifts in mistakes (and reality vs. hype) analysis he shared seems spot on in my (caveat: more limited) experience. > "IDEs/agent swarms/fallability. Both the "no need for IDE anymore" hype and the "agent swarm" hype is imo too much for right now. The models definitely still make mistakes and if you have any code you actually care about I would watch them like a hawk, in a nice large IDE on the side. The mistakes have changed a lot - they are not simple syntax errors anymore, they are subtle conceptual errors that a slightly sloppy, hasty junior dev might do. The most common category is that the models make wrong assumptions on your behalf and just run along with them without checking. They also don't manage their confusion, they don't seek clarifications, they don't surface inconsistencies, they don't present tradeoffs, they don't push back when they should, and they are still a little too sycophantic. Things get better in plan mode, but there is some need for a lightweight inline plan mode. They also really like to overcomplicate code and APIs, they bloat abstractions, they don't clean up dead code after themselves, etc. They will implement an inefficient, bloated, brittle construction over 1000 lines of code and it's up to you to be like "umm couldn't you just do this instead?" and they will be like "of course!" and immediately cut it down to 100 lines. They still sometimes change/remove comments and code they don't like or don't sufficiently understand as side effects, even if it is orthogonal to the task at hand. All of this happens despite a few simple attempts to fix it via instructions in CLAUDE . md. Despite all these issues, it is still a net huge improvement and it's very difficult to imagine going back to manual coding. TLDR everyone has their developing flow, my current is a small few CC sessions on the left in ghostty windows/tabs and an IDE on the right for viewing the code + manual edits."
- shawabawa3 8mo agoIt's been a bit like the boiling frog analogy for me I started by copy pasting more and more stuff in chatgpt. Then using more and more in-IDE prompting, then more and more agent tools (Claude etc). And suddenly I realise I barely hand code anymore For sure there's still a place for manual coding, especially schemas/queries or other fiddly things where a tiny mistake gets amplified, but the vast majority of "basic work" is now just prompting, and honestly the code quality is _better_ that it was before, all kinds of refactors I didn't think about or couldn't be bothered with have almost automatically And people still call them stochastic parrots
- phailhaus 8mo ago> And people still call them stochastic parrots Both can be true. You're tapping into every line of code publicly available, and your day-to-day really isn't that unique. They're really good at this kind of work.
- ed_mercer 8mo agoI find myself even for small work, telling CC to fix it for me is better as it usually belongs to a thread of work, and then it understands the big picture better.
- Macha 8mo agoI've had the opposite experience, it's been a long time listening to people going "It's really good now" before it developed to a permutation that was actually worth the time to use it. ChatGPT 3.5/4 (2023-2024): The chat interface was verbose and clunky and it was just... wrong... like 70+% of the time. Not worth using. CoPilot autocomplete and Gitlab Duo and Junie (late 2024-early 2025): Wayyy too aggressive at guessing exactly what I wasn't doing and hijacked my tab complete when pre-LLM type-tetris autocomplete was just more reliable. Copilot Edit/early Cursor (early 2025): Ok, I can sort of see uses here but god is picking the right files all the time such a pain as it really means I need to have figured out what I wanted to do in such detail already that what was even the point? Also the models at that time just quickly descended into incoherency after like three prompts, if it went off track good luck ever correcting it. Copilot Agent mode / Cursor (late 2025): Ok, great, if the scope is narrowly scoped, and I'm either going to write the tests for it or it's refactoring existing code it could do something. Like something mechanical like the library has a migration where we need to replace the use of methods A/B/C and replace them with a different combination of X/Y/Z. great, it can do that. Or like CRUD controller #341. I mean, sure, if my boss is going to pay for it, but not life changing. Zed Agent mode / Cursor agent mode / Claude code (early 2026): Finally something where I can like describe the architecture and requirements of a feature, let it code, review that code, give it written instructions on how to clean it up / refactor / missing tests, and iterate. But that was like 2 years of "really it's better and revolutionary now" before it actually got there. Now maybe in some languages or problem domains, it was useful for people earlier but I can understand people who don't care about "but it works now" when they're hearing it for the sixth time. And I mean, what one hand gives the other takes away. I have a decent amount of new work dealing with MRs from my coworkers where they just grabbed the requirements from a stakeholder, shoved it into Claude or Cursor and it passed the existing tests and it's shipped without much understanding. When they wrote them themselves, they tested it more and were more prepared to support it in production...
- Madmallard 8mo agoAre game developers vibe coding with agents? It's such a visual and experiential thing that writing true success criteria it can iterate on seems like borderline impossible ahead of time.
- redox99 8mo agoVibe coding in Unreal Engine is of limited use. It obviously helps with C++, but so much of your time is doing things that are not C++. It hurts a lot that UE relies heavily on blueprints, if they were code you could just vibecode a lot of that.
- 20260126032624 8mo agoI don't "vibe code" but when I use an LLM with a game I usually branch out into several experiments which I don't have to commit to. Thus, it just makes that iteration process go faster. Or slower, when the LLM doesn't understand what I want, which is a bigger issue when you spawn experiments from scratch (and have given limited context around what you are about to do).
- TheGRS 8mo agoI'm trying it out with Godot for my little side projects. It can handle writing the GUI files for nodes and settings. The workflow is asking cursor to change something, I review the code changes, then load up the game in Godot to check out the changes. Works pretty well. I'm curious if any Unity or Unreal devs are using it since I'm sure its a similar experience.
- ex-aws-dude 8mo agoA big problem is that a lot of game logic is done in visual scripting (e.g unreal blueprints) which AI tools have no idea about
- dysoco 8mo agoIt might be biased to Reddit/Twitter users but from what I've seen game developers seem to be much more averse towards using AI (even for coding) than other fields. Which is curious since prototyping helps a lot in gamedev.
- rschick 8mo agoGreat point about expansion vs speedup. I now have time to build custom tools, implement more features, try out different API designs, get 100% test coverage.. I can deliver more quickly, but can also deliver more overall.
- jopsen 8mo ago> - How much of society is bottlenecked by digital knowledge work? Any qualified guesses? I'm not convinced more traders on wall street will allocate capital more effectively leading to economic growth. Will more programmers grow the economy? Or should we get real jobs ;)
- iwontberude 8mo agoMost of this countries challenges are strictly political. The pittance of work software can contribute is most likely negligible or destructive (e.g. software buttons in cars or palantir). In other words were picked all the low hanging fruit and all that left is to hang ourselves.
- js8 8mo agoI actually disagree. Having software (AI) that can cut through the technological stuff faster will make people more aware of political problems.
- iwontberude 8mo agoedit: country's* all that is left*
- cyanydeez 8mo agoSo I'm curious, whats the actual quality control. Like, do these guys actually dog food real user experience, or are they all admins with the fast lane to the real model while everyone outside the org has to go through the 10 layers of model sheding, caching and other means and methods of saving money. We all know these models are expensive as fuck to run and these companies are degrading service, A+B testing, and the rest. Do they actually ponder these things directly? Just always seems like people are on drugs when they talk about the capabilities, and like, the drugs could be pure shit (good) or ditch weed, and we call just act like the pipeline for drugs is a consistent thing but it's really not, not at this stage where they're all burning cash through infrastructure. Definitely, like drug dealers, you know they're cutting the good stuff with low cost cached gibberish.
- bigwheels 8mo agoIf you access a model through an openrouter provider it might be quantized (akin to being "cut with trash"), but when you go directly to Anthropic or OpenAI you are getting access to the same APIs as everyone else. Even top-brass folks within Microsoft use Anthropic and OpenAI proper (not worth the red-tape trouble to go directly through Azure). Also, the creator and maintainer of Claude, Boris Cherny, was a bit of an oddball but one of the comparatively nicer people at Anthropic, and he indicated he primarily uses the same Anthropic APIs as everyone else (which makes sense from a product development perspective). The underlying models are all actually really undifferentiated under the covers except for the post-training and base prompts. If you eliminate the base prompts the models behave near identically. A conspiracy would be a helluva lot more interesting and fun, but I've spoken to these folks firsthand and it seems they already have enough challenges keeping the beast running.
- quinnjh 8mo ago> Definitely, like drug dealers, you know they're cutting the good stuff with low cost cached gibberish. Can confirm. My partner's chatGPT wouldnt return anything useful for her given a specific query involving web use, while i got the desired result sitting side by side. She contacted support and they said nothing they can do about it, her account is in an A/B test group without some features removed. I imagine this saves them considerable resources despite still billing customers for them. how much this is occurring is anyones guess
- atonse 8mo ago> LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building. I’ve always said I’m a builder even though I’ve also enjoyed programming (but for an outcome, never for the sake of the code) This perfectly sums up what I’ve been observing between people like me (builders) who are ecstatic about this new world and programmers who talk about the craft of programming, sometimes butting heads. One viewpoint isn’t necessarily more valid, just a difference of wiring.
- jimbokun 8mo agoThe new LLM centered workflow is really just a management job now. Managers and project managers are valuable roles and have important skill sets. But there's really very little connection with the role of software development that used to exist. It's a bit odd to me to include both of these roles under a single label of "builders", as they have so little in common. EDIT: this goes into more detail about how coding (and soon other kinds of knowledge work) is just a management task now: https://www.oneusefulthing.org/p/management-as-ai-superpower/ https://www.oneusefulthing.org/p/management-as-ai-superpower...
- simianwords 8mo agoi don't disagree. at some point LLM's might become good enough that we wouldn't need exact technical expertise.
- ryandrake 8mo agoI noticed the same thing, but wasn't able to put it into words before reading that. Been experimenting with LLM-based coding just so I can understand it and talk intelligently about it (instead of just being that grouchy curmudgeon), and the thought in the back of my mind while using Claude Code is always: "I got into programming because I like programming, not whatever this is..." Yes, I'm building stupid things faster, but I didn't get into programming because I wanted to build tons of things. I got into it for the thrill of defining a problem in terms of data structures and instructions a computer could understand, entering those instructions into the computer, and then watching victoriously while those instructions were executed. If I was intellectually excited about telling something to do this for me, I'd have gotten into management.
- vibeprofessor 8mo agoThe AGI vibes with Claude Code are real, but the micromanagement tax is heavy. I spend most of my time babysitting agents. I expect interviews will evolve into "build project X with an LLM while we watch" and audit of agent specs
- 0xy 8mo agoSounds great to me. Leetcode is outdated and heavily abused by people who share the questions ahead of time in various forums and chats.
- thefourthchime 8mo agoFrom what I've heard, what few interviews there are for software engineers these days, they do have you use models and see how quickly you can build things.
- iwontberude 8mo agoThe interviews I’ve given have asked about how control for AI slop without hurting your colleagues feelings. Anyone can prompt and build, the harder part, as usual for business, is knowing how and when to say, ‘no.’
- maxdo 8mo agoI've been doing vibe code interviews for nearly a year now. Most people are surprisingly bad with AI tools. We specifically ask them to bring their preferred tool, yet 20–30% still just copy-paste code from ChatGPT. fun stats: corelation is real, people who were good at vibe code, also had offer(s) with other companies that didn't run vibe code interviews.
- bflesch 8mo agoInteresting you say that, feels like when people were too stupid to google things and "googling something" was a skill that some had and others didn't.
- xyzsparetimexyz 8mo ago
- onetimeusename 8mo ago> the ratio of productivity between the mean and the max engineer? It's quite possible that this grows *a lot* I have a professor who has researched auto generated code for decades and about six months ago he told me he didn't think AI would make humans obsolete but that it was like other incremental tools over the years and it would just make good coders even better than other coders. He also said it would probably come with its share of disappointments and never be fully autonomous. Some of what he said was a critique of AI and some of it was just pointing out that it's very difficult to have perfect code/specs.
- slfreference 8mo agoI can sense two classes of coders emerging. Billionaire coder: a person who has "written" billion lines. Ordinary coders : people with only couple of thousands to their git blame.
- fishtoaster 8mo ago> if you have any code you actually care about I would watch them like a hawk, in a nice large IDE on the side. This is about where I'm at. I love pure claude code for code I don't care about, but for anything I'm working on with other people I need to audit the results - which I much prefer to do in an IDE.
- porise 8mo agoI wish the people who wrote this let us know what king of codebases they are working on. They seem mostly useless in a sufficiently large codebase especially when they are messy and interactions aren't always obvious. I don't know how much better Claude is than ChatGPT, but I can't get ChatGPT to do much useful with an existing large codebase.
- Okkef 8mo agoTry Claude code. It’s different. After you tried it, come back.
- Imustaskforhelp 8mo agoI think its not Claude code per se itself but rather the (Opus 4.5 model?) or something in an agentic workflow. I tried a website which offered the Opus model in their agentic workflow & I felt something different too I guess. Currently trying out Kimi code (using their recent kimi 2.5) for the first time buying any AI product because got it for like 1.49$ per month. It does feel a bit less powerful than claude code but I feel like monetarily its worth it. Y'know you have to like bargain with an AI model to reduce its pricing which I just felt really curious about. The psychology behind it feels fascinating because I think even as a frugal person, I already felt invested enough in the model and that became my sunk cost fallacy Shame for me personally because they use it as a hook to get people using their tool and then charge next month 19$ (I mean really Cheaper than claude code for the most part but still comparative to 1.49$)
- maxdo 8mo agochatGPT is not made to write code. Get out of stone age :)
- CameronBanga 8mo agoThis is an antidotal example, but I released this last week after 3 months of work on it as a "nights and weekdends" project: https://apps.apple.com/us/app/skyscraper-for-bluesky/id6754198379 https://apps.apple.com/us/app/skyscraper-for-bluesky/id67541... I've been working in the mobile space since 2009, though primarily as a designer and then product manager. I work in kinda a hybrid engineering/PM job now, and have never been a particularly strong programmer. I definitely wouldn't have thought I could make something with that polish, let alone in 3 months. That code base is ~98% Claude code.
- strogonoff 8mo agoLLM coding splits up engineers based on those who primarily like building and those who primarily like code reviews and quality assessment. I definitely don’t love the latter (especially when reviewing decisions not made by a human with whom I can build long-term personal rapport). After certain experience threshold of making things from scratch, “coding” (never particularly liked that term) has always been 99% building, or architecture, and I struggle to see how often a well-architected solution today, with modern high-level abstractions, requires so much code that you’d save significant time and effort by not having to just type, possibly with basic deterministic autocomplete, exactly what you mean (especially considering you would have to also spend time and effort reviewing whatever was typed for you if you used a non-deterministic autocomplete).
- OkayPhysicist 8mo agoSee, I don't take it that extreme: LLMs make fantastic, never-before seen quality autocompletes. I hacked together a Neovim plugin that prompts an LLM to "finish this function" on command, and it's a big time save for the menial plumbing type operations. Think things like "this api I use expects JSON that encodes some subset of SQL, I want all the dogs with Ls in their name that were born on a Tuesday". Given an example of such API (or if the documentation ended up in its training), LLMs will consistently one-shot stuff like that. Asking it to do entire projects? Dumb. You end up with spaghetti, unless you hand-hold it to a point that you might as well be using my autocomplete method.
- gverrilla 8mo agoDepends on the scope of the project. If it's small, and you direct it correctly, it can one-shot yes. Or 2-3-shot.
- cmrdporcupine 8mo ago"those who primarily like code reviews and quality assessment" -- I don't love those. In fact I find it tedious and love it when I can work on my own without them. Except after 25 years of working I know how imperative they are, how easily a project can disintegrate into confused silos, and am frustrated as heck with these tools being pushed without attention to this problem.
- DeathArrow 8mo ago>LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building. Quite insightful.
- jimbokun 8mo agoI'm pretty happy with Copilot in VS Code. Type what change I want Claude to make in the Copilot panel, and then use the VS Code in context diffs to accept or reject the proposed changes. While being able to make other small changes on my own. So I think this tracks with Karpathy's defense of IDEs still being necessary ? Has anyone found it practical to forgo IDEs almost entirely?
- maxdo 8mo agoCoplilot is not on par with cc or cursor even
- WA 8mo agoWhy not? You can select Opus 4.5, Gemini 3 Pro, and others.
- spaceman_2020 8mo agoClaude Code is a CLI tool which means it can do complete projects in a single command. Also has fantastic tools for scaffolding and harnessing the code. You can define everything from your coding style to specific instructions for designing frontpages, integrating payments, etc. It's not about the model. It's about the harness
- piker 8mo agoThis would make some sense if VS Code didn't have a terminal built into it. The LLMs have the same bash capabilities in either form.
- binarycrusader 8mo agoClaude Code is a CLI tool which means it can do complete projects in a single command https://github.com/features/copilot/cli/ https://github.com/features/copilot/cli/
- sandos 8mo agoHuh? There is nothing stopping copilot from doing an entire project in one go. Ive done it 10s of times.
- uejfiweun 8mo agoHonestly, how long do you guys think we have left as SWEs with high pay? Like the SWE job will still exist, but with a much lower technical barrier of entry, it strikes me that the pay is going to decrease a lot. Obviously BigCo codebases are extremely complex, more than Claude Code can handle right now, but I'd say there's definitely a timer running here. The big question for my life personally is whether I can reach certain financial milestones before my earnings potential permanently decreases.
- spaceman_2020 8mo agoI think the senior devs will be fine. They're like lawyers at this point - everyone is too scared they'll screw up and will keep them around The juniors though will radically have to upskill. The standard junior dev portfolio can be replicated by claude code in like three prompts The game has changed and I don't think all the players are ready to handle it
- jerf 8mo agoIt's counterintuitive but something becoming easier doesn't necessarily mean it becomes cheap. Programming has arguably been the easiest engineering discipline to break into by sheer force of will for the past 20+ years, and the pay scales you see are adapted to that reality already. Empowering people to do 10 times as much as they could before means they hit 100 times the roadblocks. Again, in a lot of ways we've already lived in that reality for the past many years. On a task-by-task basis programming today is already a lot easier than it was 20 years ago, and we just grew our desires and the amount of controls and process we apply. Problems arise faster than solutions. Growing our velocity means we're going to hit a lot more problems. I'm not saying you're wrong, so much as saying, it's not the whole story and the only possibility. A lot of people today are kept out of programming just because they don't want to do that much on a computer all day, for instance. That isn't going to change. There's still going to be skills involved in being better than other people at getting the computers to do what you want. Also on a long term basis we may find that while we can produce entry-level coders that are basically just proxies to the AI by the bucketful that it may become very difficult to advance in skills beyond that, and those who are already over the hurdle of having been forced to learn the hard way may end up with a very difficult to overcome moat around their skills, especially if the AIs plateau for any period of time. I am concerned that we are pulling up the ladder in a way the ladder has never been pulled up before.
- hollowturtle 8mo ago> Coding workflow. Given the latest lift in LLM coding capability, like many others I rapidly went from about 80% manual+autocomplete coding and 20% agents in November to 80% agent coding and 20% edits+touchups in December Anyone wondering what exactly is he actually building? What? Where? > The mistakes have changed a lot - they are not simple syntax errors anymore, they are subtle conceptual errors that a slightly sloppy, hasty junior dev might do. I would LOVE to have jsut syntax errors produced by LLMs, "subtle conceptual errors that a slightly sloppy, hasty junior dev might do." are neither subtle nor slightly sloppy, they actually are serious and harmful, and no junior devs have no experience to fix those. > They will implement an inefficient, bloated, brittle construction over 1000 lines of code and it's up to you to be like "umm couldn't you just do this instead?" Why just not hand write 100 loc with the help of an LLM for tests, documentation and some autocomplete instead of making it write 1000 loc and then clean it up? Also very difficult to do, 1000 lines is a lot. > Tenacity. It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a computer program running in the cloud, what exactly did he expected? > Speedups. It's not clear how to measure the "speedup" of LLM assistance. See above > 2) I can approach code that I couldn't work on before because of knowledge/skill issue. So certainly it's speedup, but it's possibly a lot more an expansion. mmm not sure, if you don't have domain knowledge you could have an initial stubb at the problem, what when you need to iterate over it? You don't if you don't have domain knowledge on your own > Fun. I didn't anticipate that with agents programming feels more fun because a lot of the fill in the blanks drudgery is removed and what remains is the creative part. No it's not fun, eg LLMs produce uninteresting uis, mostly bloated with react/html > Atrophy. I've already noticed that I am slowly starting to atrophy my ability to write code manually. My bet is that sooner or later he will get back to coding by hand for periods of time to avoid that, like many others, the damage overreliance on these tools bring is serious. > Largely due to all the little mostly syntactic details involved in programming, you can review code just fine even if you struggle to write it. No programming it's not "syntactic details" the practice of programming it's everything but "syntactic details", one should learn how to program not the language X or Y > What happens to the "10X engineer" - the ratio of productivity between the mean and the max engineer? It's quite possible that this grows a lot. Yet no measurable econimic effects so far > Armed with LLMs, do generalists increasingly outperform specialists? LLMs are a lot better at fill in the blanks (the micro) than grand strategy (the macro). Did people with a smartphone outperformed photographers?
- maximedupre 8mo ago> It hurts the ego a bit but the power to operate over software in large "code actions" is just too net useful It does hurt, that's why all programmers now need an entrepreneurial mindset... you become if you use your skills + new AI power to build a business.
- xyzsparetimexyz 8mo agoWhat about the people who dont want to be entrepreneurs?
- maximedupre 8mo agoThey have to pivot to something else
- maximedupre 8mo agoOr stay ahead of the curve as long as possible, e.g. work on the loop/ralphing
- webdevver 8mo agopermanent underclass...
- jetsetk 8mo agoThat is motivational content, but not economics. Most startups will be noise, even more so than before. The value of being a founder ceases when everyone is a founder, when it becomes universal. You will need customers. Nobody wants to buy re-invented-the-wheel-74.0. It lacks character, it lacks soul. Without it, your product will be nothing but noise in a noisy world.
- maximedupre 8mo agoCope. If you create something that genuinely solves a problem, people will buy no matter what. Look entrepreneurship has never been easy. In fact it's always been one of the hardest thing ever. I'm just saying... *you don't have to do it*. Do whatever you want lol Happy to hear what's your solution to avoid becoming totally replaceable and obsolete.
- spaceman_2020 8mo agoOnce again, 80% of the comments here are from boomers. HN used to be a proper place for people actually curious about technology
- weirdmantis69 8mo agoYa it's so weird lol
- vardalab 8mo agoI'm almost a boomer and I agree. THis dichotomy is weird. I am retired EE and I love the ability to just have AI do whatever I want for me. I have it manage a 10 node proxmox cluster in my basement via ansible and terraform. I can finally do stuff I always wanted but had no time. I got sick of editing my kids sports videos for highlights in Davinci Resolve so just asked claude to write a simple app for me and then use all my random video cards in my boxes to render clips in parallel and so on. Tech is finally fun again when I do not have to dedicate days to understand some new framework. It does feel a little like late 1990's computing when everyone was making geocities webpages but those days were more fun. Now with local llms getting strong as well and speaking to my PC instead of typing it feels like SciFi, so yeah, I do not get this hacker news hand wringing about code craft.
- kejaed 8mo agoSo what is your workflow now with this app for kids sports highlights?
- zennit 8mo agoAlso interested
- vardalab 8mo agoWell, it's not really a full-blown app yet. Claude wrote a plugin for MPV. So now when I watch video I just push a button to mark in and out of highlights similar to how it works in DaVinci Resolve. Then I have a command line tool that takes those timestamps in a video file and cuts it up into individual clips and then re-renders those clips and creates a highlight reel. Another command line tool takes three or four large MP4 files that the camera generates and downloads them and combines them in the actual game video on my desktop and also uploads it to my archive and transcodes into a bunch of different formats and uploads to YouTube. And for transcoding, again, it divvies it out to the video cards, which works pretty well. I think I have five or six encoders available so it chunks it up and then reassembles. All in all, it's nothing fancy, but it reduced quite a bit the friction of coming home after games and getting a video up on YouTube for grandparents.
- einrealist 8mo ago> It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. Somewhere, there are GPUs/NPUs running hot. You send all the necessary data, including information that you would never otherwise share. And you most likely do not pay the actual costs. It might become cheaper or it might not, because reasoning is a sticking plaster on the accuracy problem. You and your business become dependent on this major gatekeeper. It may seem like a good trade-off today. However, the personal, professional, political and societal issues will become increasingly difficult to overlook.
- daxfohl 8mo agoI still find in these instances there's at least a 50% chance it has taken a shortcut somewhere: created a new, bigger bug in something that just happened not to have a unit test covering it, or broke an "implicit" requirement that was so obvious to any reasonable human that nobody thought to document it. These can be subtle because you're not looking for them, because no human would ever think to do such a thing. Then even if you do catch it, AI: "ah, now I see exactly the problem. just insert a few more coins and I'll fix it for real this time, I promise!"
- gtowey 8mo agoThe value extortion plan writes itself. How long before someone pitches the idea that the models explicitly almost keep solving your problem to get you to keep spending? Would you even know?
- fragmede 8mo agoThe free market proposition is that competition (especially with Chinese labs and grok) means that Anthropic is welcome to do that. They're even welcome to illegally collude with OpenAi such that ChatGPT is similarly gimped. But switching costs are pretty low. If it turns out I can one shot an issue with Qwen or Deepseek or Kimi thinking, Anthropic loses not just my monthly subscription, but everyone else's I show that too. So no, I think that's some grade A conspiracy theory nonsense you've got there.
- rileymichael 8mo ago> LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building as the former, i've never felt _more ahead_ than now due to all of the latter succumbing to the llm hype
- daxfohl 8mo agoI worry about the "brain atrophy" part, as I've felt this too. And not just atrophy, but even moreso I think it's evolving into "complacency". Like there have been multiple times now where I wanted the code to look a certain way, but it kept pulling back to the way it wanted to do things. Like if I had stated certain design goals recently it would adhere to them, but after a few iterations it would forget again and go back to its original approach, or mix the two, or whatever. Eventually it was easier just to quit fighting it and let it do things the way it wanted. What I've seen is that after the initial dopamine rush of being able to do things that would have taken much longer manually, a few iterations of this kind of interaction has slowly led to a disillusionment of the whole project, as AI keeps pushing it in a direction I didn't want. I think this is especially true if you're trying to experiment with new approaches to things. LLMs are, by definition, biased by what was in their training data. You can shock them out of it momentarily, whish is awesome for a few rounds, but over time the gravitational pull of what's already in their latent space becomes inescapable. (I picture it as working like a giant Sierpinski triangle). I want to say the end result is very akin to doom scrolling. Doom tabbing? It's like, yeah I could be more creative with just a tad more effort, but the AI is already running and the bar to seeing what the AI will do next is so low, so....
- Imustaskforhelp 8mo ago> I want to say it's very akin to doom scrolling. Doom tabbing? It's like, yeah I could be more creative with just a tad more effort, but the AI is already running and the bar to seeing what the AI will do next is so low, so.... Yea exactly, Like we are just waiting so that it gets completed and after it gets completed then what? We ask it to do new things again. Just as how if we are doom scrolling, we watch something for a minute then scroll down and watch something new again. The whole notion of progress feels completely fake with this. Somehow I guess I was in a bubble of time where I had always end up using AI in web browsers (just as when chatgpt 3 came) and my workflow didn't change because it was free but recently changed it when some new free services dropped. "Doom-tabbing" or complete out of the loop AI agentic programming just feels really weird to me sucking the joy & I wouldn't even consider myself a guy particular interested in writing code as I had been using AI to write code for a long time. I think the problem for me was that I always considered myself a computer tinker before coder. So when AI came for coding, my tinkering skills were given a boost (I could make projects of curiosity I couldn't earlier) but now with AI agents in this autonomous esque way, it has come for my tinkering & I do feel replaced or just feel like my ability of tinkering and my interests and my knowledge and my experience is just not taken up into account if AI agent will write the whole code in multi file structure, run commands and then deploy it straight to a website. I mean my point is tinkering was an active hobby, now its becoming a passive hobby, doom-tinkering? I feel like I have caught up on the feeling a bit earlier with just vibe from my heart but is it just me who feels this or? What could be a name for what I feel?
- TheGRS 8mo agoI do feel a big mood shift after late November. I switched to using Cursor and Gemini primarily and it was big change in my ability to get my ideas into code effectively. The Cursor interface for one got to a place that I really like and enjoy using, but its probably more that the results from the agents themselves are less frustrating. I can deal with the output more now. I'm still a little iffy on the agent swarm idea. I think I will need to see it in action in an interface that works for me. To me it feels like we are anthropomorphizing agents too much, and that results in this idea that we can put agents into roles and them combine them into useful teams. I can't help seeing all agents as the same automatons and I have trouble understanding why giving an agent with different guideliens to follow, and then having them follow along another agent would give me better results than just fixing the context in the first place. Either that or just working more on the code pipeline to spot issues early on - all the stuff we already test for.
- Macha 8mo ago> - What does LLM coding feel like in the future? Is it like playing StarCraft? Playing Factorio? Playing music? Starcraft and Factorio are exactly what it is not. Starcraft has a loooot of micro involved at any level beyond mid level play, despite all the "pro macros and beats gold league with mass queens" meme videos. I guess it could be like Factorio if you're playing it by plugging together blueprint books from other people but I don't think that's how most people play. At that level of abstraction, it's more like grand strategy if you're to compare it to any video game? You're controlling high level pushes and then the units "do stuff" and then you react to the results.
- deleted 8mo ago[deleted]
- zetazzed 8mo agoIt's like the Victoria 3 combat system. You just send an army and a general to a given front and let them get to work with no micro. Easy! But of course some percentage of the time they do something crazy like deciding to redeploy from your existential Franco-Prussian war front to a minor colonial uprising...
- kridsdale3 8mo agoI think the StarCraft analogy is fine, you have to compare it not to macro and micro RTS play, but to INDIVIDUAL UNITS. For your whole career until now, you have been a single Zergling or Probe. Now you are the Commander.
- TheRoque 8mo agoExcept that pro starcraft player still micro-manage every single Zergling or probe when necessary, while vibe coders just right click on the ennemy base and hope it'll go well
- daxfohl 8mo agoI'm curious to see what effect this change has on leadership. For the last two years it's been "put everything you can into AI coding, or else!" with quotas and firings and whatever else. Now that AI is at the stage where it can actually output whole features with minimal handholding, is there going to be a Frankenstein moment where leadership realizes they now have a product whose codebase is running away from their engineering team's ability to support it? Does it change the calculus of what it means to be underinvested vs overinvested in AI, and what are the implications?
- philipwhiuk 8mo ago> It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. The bits left unsaid: 1. Burning tokens, which we charge you for 2. My CPU does this when I tell it to do bogosort on a million 32-bit integers, it doesn't mean it's a good thing
- tintor 8mo ago"you can review code just fine even if you struggle to write it." Well, merely approving code takes no skill at all.
- roblh 8mo agoSeriously, that’s a completely nonsense line.
- forrestthewoods 8mo agoHN should ban any discussion on “things I learned playing with AI” that don’t include direct artifacts of the thing built. We’re about a year deep into “AI is changing everything” and I don’t see 10x software quality or output. Now don’t get me wrong I’m a big fan of AI tooling and think it does meaningfully increase value. But I’m damn tired of all the talk with literally nothing to show for it or back it up.
- lomase 8mo ago[dead]
- twa927 8mo agoI don't see the AI capacity jump in the recent months at all. For me it's more the opposite, CC works worse than a few months ago. Keeps forgetting the rules from CLAUDE.md, hallucinates function calls, generates tons of over-verbose plans, generates overengineered code. Where I find it a clear net-positive is pure frontend code (HTML + Tailwind), it's spaghetti but since it's just visualization, it's OK.
- ValentineC 8mo ago> Where I find it a clear net-positive is pure frontend code (HTML + Tailwind), it's spaghetti but since it's just visualization, it's OK. This makes it sound like we're back in the days of FrontPage/Dreamweaver WYSIWYG. Goodness.
- twa927 8mo agoHmm, your comment gave me the idea that maybe we should invent "What You Describe Is What You Get|. To replace HTML+Tailwind spaghetti with prompts generating it.
- DominikPeters 8mo agoAre you using Opus 4.5? Sounds more like Sonnet.
- twa927 8mo agoYes I'm using Sonnet 4.5. Thanks for the tip, will try Opus 4.5, although costs might become an issue.
- TuxSH 8mo ago> although costs might become an issue. If you have a ChatGPT subscription, try Codex with GPT-5.2-High or 5.2-codex High? In my experience, while being much slower, it produces far better results than Opus and seems even more aggressively subsidized (more generous rate limits).
- 8mo ago
- nsb1 8mo agoThe best thing I ever told Claude to do was "Swear profusely when discussing code and code changes". Probably says more about me than Claude, but it makes me snicker.
- neuralkoi 8mo ago> The most common category is that the models make wrong assumptions on your behalf and just run along with them without checking. If current LLMs are ever deployed in systems harboring the big red button, they WILL most definitely somehow press that button.
- arthurcolle 8mo agoUS MIC are already planning on integrating fucking Grok into military systems. No comment.
- groby_b 8mo agofwiw, the same is true for humans. Which is why there's a whole lot of process and red tape around that button. We know how to manage risk. We can choose to do that for LLM usage, too. If instead we believe in fantasies of a single all-knowing machine god that is 100% correct at all times, then... we really just have ourselves to blame. Might as well just have spammed that button by hand.
- 0xbadcafebee 8mo ago> What happens to the "10X engineer" - the ratio of productivity between the mean and the max engineer? It's quite possible that this grows a lot. I was thinking about this the other day as relates to the DevOps movement. The DevOps movement started as a way to accelerate and improve the results of dev<->ops team dynamics. By changing practices and methods, you get acceleration and improvement. That creates "high-performing teams", which is the team form of a 10x engineer. Whether or not you believe in '10x engineers', a high-performing team is real. You really can make your team deploy faster, with fewer bugs. You have to change how you all work to accomplish it, though. To get good at using AI for coding, you have to do the same thing: continuous improvement, changing workflows, different designs, development of trust through automation and validation. Just like DevOps, this requires learning brand new concepts, and changing how a whole team works. This didn't get adopted widely with DevOps because nobody wanted to learn new things or change how they work. So it's possible people won't adapt to the "better" way of using AI for coding, even if it would produce a 10x result. If we want this new way of working to stick, it's going to require education, and a change of engineering culture.
- virgilp 8mo agoThis is an interesting thing that I'm contemplating. I also do believe that (perhaps with very few exceptions) there are no "10x engineers" by themselves, but engineers that thrive 10x more in a context or another (like, I'm sure Jeff Dean is an absolutely awesome engineer - but if you took him out of Google and plugged him into IBM - would he have had the same impact?) With that in mind - I think one very unexplored area is "how to make the mixed AI-human teams successful". Like, I'm fairly convinced AI changes things, but to get to the industrialization of our craft (which is what management seems to want - and, TBH, something that makes sense from an economic pov), I feel that some big changes need to happen, and nobody is talking about that too much. What are the changes that need to happen? How do we change things, if we are to attempt such industrialization?
- superze 8mo agoI don't know about you guys but most of the time it's spitting nonsense models in sqlalchemy and I have to constantly correct it to the point where I am back at writing the code myself. The bugs are just astonishing and I lose control of the codebase after some time to the point where reviewing the whole thing just takes a lot of time. On the contrary if it was for a job in a public sector I would just let the LLM spit out some output and play stupid, since salary is very low.
- all2well 8mo agoWhat particular setups are getting folks these sorts of results? If there’s a way I could avoid all the babysitting I have to do with AI tools that would be welcome
- spongebobstoes 8mo agoi use codex cli. work on giving it useful skills. work on the other instruction files. take Karpathy tips around testing and declarativeness use many simultaneously, and bounce between them to unblock them as needed build good tools and tests. you will soon learn all the things you did manually -- script them all
- geraneum 8mo ago> If there’s a way I could avoid all the babysitting I have to do with AI tools that would be welcome OP mentions that they are actually doing the “babysitting”
- toephu2 8mo agoI think in less than a year writing code manually will be akin to doing arithmetic problems by hand. Sure you can still code manually, but it's going to be a lot faster to use an LLM (calculator).
- kypro 8mo agoI agree, but writing code is so different to calculations that long-term benefits are less clear. It doesn't matter how good you are at calculations the answer to 2 + 2 is always 4. There are no methods of solving 2 + 2 which could result in you accidentally giving everyone who reads the result of your calculation write access to your entire DB. But there are different ways to code a system even if the UI is the same, and some of these may neglect to consider permissions. I think a good parallel here would be to imagine that tomorrow we had access to humanoid robots who could do construction work. Would we want them to just go build skyscrapers and bridges and view all construction businesses which didn't embrace the humanoid robots as akin to doing arithmetic by hand? You could of course argue that there's no problem here so long as trained construction workers are supervising the robots to make sure they're getting tolerances right and doing good welds, but then what happens 10 years down the road when humans haven't built a building in years? If people are not writing code any more then how can people be expected to review AI generated code? I think the optimistic picture here is that humans just won't be needed in the future. In theory when models are good enough we should be able to trust the AI systems more than humans. But the less optimistic side of me questions a future in which humans no longer do, or even know how to do such fundamental things.
- adamddev1 8mo agoPeople keep using these analogies but I think these are fundamentally different things. 1. hand arithmetic -> using a calculator 2. assembly -> using a high level language 3. writing code -> making an LLM write code Number 3 does not belong. Number 3 is a fundamentally different leap because it's not based on deterministic logic. You can't depend on an LLM like you can depend on a calculator or a compiler. LLMs are totally different.
- Havoc 8mo ago
- nsainsbury 8mo agoTouching on the atrophy point, I actually wrote a few thoughts about this yesterday: https://www.neilwithdata.com/outsourced-thinking https://www.neilwithdata.com/outsourced-thinking I actually disagree with Andrej here re: "Generation (writing code) and discrimination (reading code) are different capabilities in the brain." and I would argue that the only reason he can read code fluently, find issues, etc. is because he has spent year in a non-AI assisted world writing code. As time goes on, he will become substantially worse. This also bodes incredibly poorly for the next generation, who will mostly in their formative years now avoid writing code and thus fail to even develop a idea of what good code is, how it works/why it works, why you make certain decisions, and not others, etc. and ultimately you will see them become utterly dependent on AI, unable to make progress without it. IMO outsourcing thinking is going to have incredibly negative consequences for the world at large.
- gwd 8mo agoIs coding like piloting, where pilots need a certain number of hours of "flight time" to gain skills, and then a certain number of additional hours each year to maintain their skills? Do developers need to schedule in a certain number of "manually written lines of code" every year?
- thoughtpeddler 8mo agoRead your blog post and agree with some of it. Largely I agree with the premise that the 2nd and 3rd order effects of this technology will be more impactful than the 1st order “I was able to code this app I wouldn’t have otherwise even attempted to”. But they are so hard to predict!
- olafalo 8mo agoThanks, this rings true to me. The struggle is an investment, and it pays off in good judgement and taste. The same goes for individual codebases too. When I see some weird bug and can immediately guess what’s going wrong and why, that’s my time spent in that codebase paying off. I guess LLM-ing a feature is the inverse, incurring some kind of cognitive debt.
- nicodjimenez 8mo ago
- themafia 8mo agoInstead of a 17 paragraph twitter post with a baffling TLDR at the end why not just record your screen and _demonstrate_ all of what you're describing? Otherwise, I think you're incidentally right, your "ego" /is/ bruised, and you're looking for a way out by trying to prognosticate on the future of the technology. You're failing in two different ways.
- deleted 8mo ago[deleted]
- alexose 8mo agoIt's refreshing to see one of the top minds in AI converge on the same set of thoughts and frustrations as me. For as fast as this is all moving, it's good to remember that most of us are actually a lot closer to the tip of the spear than we think.
- siliconc0w 8mo agoNot sure how he is measuring, I'm still closer to about a 60% success rate. It's more like 20% is an acceptable one-shot, this goes to 60% acceptable with some iteration, but 40% either needs manual intervention to succeed or such significant iteration that manual is likely faster. I can supervise maybe three agents in parallel before a task requiring significant hand-holding means I'm likely blocking an agent. And the time an agent is 'restlessly working' on something in usually inversely correlated with the likelihood to succeed. Usually if it's going down a rabbit hole, the correct thing to do is to intervene and reorient it.
- randoglando 8mo agoSenpai has taken the words out of my mouth and put them on the page.
- ositowang 8mo agoIt’s a great and insightful review—not over-hyping the coding agent, and not underestimating it either. It acknowledges both its usefulness and its limitations. Embracing it and growing with it is how I see it too.
- epolanski 8mo ago> What happens to the "10X engineer" - the ratio of productivity between the mean and the max engineer? It's quite possible that this grows a lot. No doubt that good engineers will know when and how to leverage the tool, both for coding and improving processes (design-to-code, requirement collection, task tracking, basic code reviewal, etc) improving their own productivity and of those around them. Motivated individuals will also leverage these tools to learn more and faster. And yes, of course it's not the only tool one should use, of course there's still value in talking with proper human experts to learn from, etc, but 90% of the time you're looking for info the LLM will dig it from you reading at the source code of e.g. Postgres and its test rather than asking on chats/stack overflow. This is a trasformative technology that will make great engineers even stronger, but it will weed out those who were merely valued for their very basic capability of churning something but never cared neither about engineering nor coding, which is 90% of our industry.
- appstorelottery 8mo ago> Atrophy. I've already noticed that I am slowly starting to atrophy my ability to write code manually. I've been increasingly using LLM's to code for nearly two years now - and I can definitely notice my brain atrophy. It bothers me. Actually over the last few weeks I've been looking at a major update to a product in production & considered doing the edits manually - at least typing the code from the LLM & also being much more granular with my instructions (i.e. focus on one function at a time). I feel in some ways like my brain is turning into slop & I've been coding for at least 35 years... I feel validated by Karpathy.
- epolanski 8mo agoDon't be too worried about it. 1. Manual coding may be less relevant (albeit ability to read code, interpret it and understand it will be more) in the future. Likely already is. 2. Any skill you don't practice becomes "weaker". Gonna give you an example. I play chess since my childhood, but sometimes I go months without playing it, even years. When I get back I start losing elo fast. If I was in the top 10% of chess.com, I drop to top 30% in the weeks after. But after few months I'm back at top 10%. Takeaway: your relative ability is more or less the same compared to other practitioners, you're simply rusty.
- appstorelottery 8mo agoThanks for your comment, it set me at ease. I know from experience that you're right on point 2. As for point one, I also tend to agree. AI is such a paradigm shift & rapid/massive change doesn't come without stress. I just need to stay cool about it all ;-)
- oxag3n 8mo ago> Atrophy. I've already noticed that I am slowly starting to atrophy my ability to write code manually... > Largely due to all the little mostly syntactic details involved in programming, you can review code just fine even if you struggle to write it. Until you struggle to review it as well. Simple exercise to prove it - ask LLM to write a function in familiar programming language, but in the area you didn't invest learning and coding yourself. Try reviewing some code involving embedding/SIMD/FPGA without learning it first.
- sleazebreeze 8mo agoPeople would struggle to review code in a completely unfamiliar domain or part of the stack even before LLMs.
- chrisjj 8mo agoNo, because they wouldn't be so foolish as to try it.
- piskov 8mo agoThat’s why you need to write code to learn it. No-one has ever learned skill just by reading/observing
- sponaugle 8mo ago"No-one has ever learned skill just by reading/observing" - Except of course all of those people in Cosmology who, you know, observe.
- direwolf20 8mo agowhat skill do they have? making stars? no they are skilled at observing, which is what they do.
- sponaugle 8mo ago
- thomassmith65 8mo agoSlopacolypse. I am bracing for 2026 as the year of the slopacolypse across all of github, substack, arxiv, X/instagram, and generally all digital media. Did he coin the term "slopacolypse"? It's a useful one.
- chrisjj 8mo agoI prefer slopocalypse.
- rvz 8mo agoThat works better. “slopacolypse” does not make any sense both in writing and pronunciation.
- direwolf20 8mo agonot aslopalypse?
- direwolf20 8mo agoor even aislopalypse?
- pron 8mo agoPeople who just let the agent code for them, how big of a codebase are you working on? How complex (i.e. is it a codebase that junior programmers could write and maintain)?
- aixpert 8mo agorust compiler and redox operating system with modified Qemu for Mac Vulcan metal pipeline ... probably not junior stuff you might think I'm kidding but Search redox on github, you will find that project and the anonymous contributions
- rester324 8mo agoI am curious. What do you want us to see in that github repo?
- bojo 8mo agoI've been an EM for the last 10 of my 25 year Software Engineering career. Coding is, frankly, boring to me anymore, even though I enjoyed doing it most of my career. I had this project I wanted to exist in world but couldn't be bothered to get started. Decided to figure out what this "vibe coding" nonsense is, and now there's a certain level of joy to all of this again. Being able to clearly define everything using markdown contexts before any code is even written has been a great way to brain dump those 25 years of experience and actually watch something sane get produced. Here are the stats Claude Code gave me: Overview ┌───────────────┬────────────────────────────┐ │ Metric │ Value │ ├───────────────┼────────────────────────────┤ │ Total Commits │ 365 │ ├───────────────┼────────────────────────────┤ │ Project Age │ 7 days (Jan 20 - 27, 2026) │ ├───────────────┼────────────────────────────┤ │ Open Issues │ 5 │ ├───────────────┼────────────────────────────┤ │ Contributors │ 1 │ └───────────────┴────────────────────────────┘ Lines of Code by Language ┌───────────────────────────┬───────┬────────┬───────────┐ │ Language │ Files │ Lines │ % of Code │ ├───────────────────────────┼───────┼────────┼───────────┤ │ Rust (Backend) │ 94 │ 31,317 │ 51.8% │ ├───────────────────────────┼───────┼────────┼───────────┤ │ TypeScript/TSX (Frontend) │ 189 │ 29,167 │ 48.2% │ ├───────────────────────────┼───────┼────────┼───────────┤ │ SQL (Migrations) │ 34 │ 1,334 │ — │ ├───────────────────────────┼───────┼────────┼───────────┤ │ CSS │ — │ 1,868 │ — │ ├───────────────────────────┼───────┼────────┼───────────┤ │ Markdown (Docs) │ 37 │ 9,485 │ — │ ├───────────────────────────┼───────┼────────┼───────────┤ │ Total Source │ 317 │ 60,484 │ 100% │ └───────────────────────────┴───────┴────────┴───────────┘
- tomlockwood 8mo agoOh wow! Guy who's current project depends on AI being good is talking about AI being good. Interesting.
- bartoszcki 8mo agoFeels like a combination of writing very detailed task descriptions and reviewing junior devs. It's horrible. I very much hope this won't be my job.
- giancarlostoro 8mo ago> IDEs/agent swarms/fallability. Both the "no need for IDE anymore" hype and the "agent swarm" hype is imo too much for right now. I'm honestly considering throwing away my JetBrains subscription and this is year 9 or 10 of me having one. I only open Zed and start yappin' at Claude Code. My employer doesn't even want me using ReSharper because some contractor ruined it for everyone else by auto running all code suggestions and checking them in blindly, making for really obnoxious code diffs and probably introducing countless bugs and issues. Meanwhile tasks that I know would take any developers months, I can hand-craft with Claude in a few hours, with the same level of detail, but no endless weeks of working on things that'll be done SoonTM.
- jedberg 8mo ago> You realize that stamina is a core bottleneck to work There has been a lot of research that shows that grit is far more correlated to success than intelligence. This is an interesting way to show something similar. AIs have endless grit (or at least as endless as your budget). They may outperform us simply because they don't ever get tired and give up. Full quote for context: Tenacity. It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. You realize that stamina is a core bottleneck to work and that with LLMs in hand it has been dramatically increased.
- Loeffelmann 8mo agoIf you ever work with LLMs you know that they quite frequently give up. Sometimes it's a // TODO: implement logic or a "this feature would require extensive logic and changes to the existing codebase". Sometimes they just declare their work done. Ignoring failing tests and builds. You can nudge them to keep going but I often feel like, when they behave like this, they are at their limit of what they can achieve.
- energy123 8mo agoUsing LLMs to clean those up is part of the workflow that you're responsible for (... for now). If you're hoping to get ideal results in a single inference, forget it.
- jedberg 8mo ago> If you ever work with LLMs you know that they quite frequently give up. If you try to single shot something perhaps. But with multiple shots, or an agent swarm where one agent tells another to try again, it'll keep going until it has a working solution.
- alansaber 8mo agoYeah exactly this is a scope problem, actual input/output size is always limited> I am 100% sure CC etc are using multiple LLM calls for each response, even though from the response streaming it looks like just one.
- ares623 8mo agoImagine taking career advice from people who will never need to be employed again in order to survive.
- fragmede 8mo agoYes, typically you take since from people who've been successful at their career. Are you suggesting we should be taking career advice from high school freshmen instead?
- ares623 8mo agoI'm nitpicking on the atrophy bit. He can afford to have his skills or his brain atrophied. His followers though? Nevermind the fact he became successful _because_ of his skills and his brain.
- longhaul 8mo agoAm working on an iPhone app and impressed with how well Claude is able to generate decent/working code with prompts in plain English. I don’t have previous experience in building apps or swift but have a C++ background. Working in smaller chunks and incrementally adding features rather than a large prompt for the whole app seems more practical, is easier to review and build confidence. Adding/prompting features one by one, reviewing code and then testing the resulting binary feels like the new programming workflow Prompt/REview/Test - PRET.
- wellpast 8mo agoI coded up a crossword puzzle game using agentic dev this weekend. Claude and Codex/GPT. Had to seriously babysit and rewrite much of it, though, sure, I found it “cool” what it could do. Writing code in many cases is faster to me than writing English (that is how PLs are designed, btw!) LLM/agentic is very “neat” but still a toy to the professional, I would say. I doubt reports like this one. For those of us building real world products with shelf-lives (Is Andrej representative of this archetype?), I just don’t see the value-add touted out there. I’d love to be proven wrong. But writing code (in code, not English), to me and many others, is still faster than reading/proving it. I think there’s a combination of fetishizing and Stockholm syndroming going on in these enthusiastic self-reports. PMW.
- jofla_net 8mo ago>Writing code in many cases is faster to me than writing English True, I feel as though i'd have to become Stienbeck to get it to do what i "really" wanted, with all the true nuance.
- axus 8mo agoFinally, literate programming! https://en.wikipedia.org/wiki/Literate_programming https://en.wikipedia.org/wiki/Literate_programming
- lofaszvanitt 8mo agoThe whole thing is about getting rid of experts and let the entry level idiots do all the work. The coders become expendable. And people do not see the chasm staring back at them :D. LLMs in their current form redistributes "intelligence" and expertise to the average joes for mere pennies. It should be much much more expensive, or it will disrupt the whole ecosystem. If it becomes even more intelligent it must be bludgeoned to death a.k.a. regulated like hell, otherwise the ensuing disruption will kill the job market and in the long term human values. As an added plus: those, who already have wealth will benefit the most, instead of the masses. Since the distribution and dissemination of new projects is at the same level as before, meaning you would need a lot of money. So no matter how clever you are with an llm, if you don't have the means to distribute it you will be left in the dirt.
- kshri24 8mo agoAgree with Karpathy's take. Finally a down to Earth analysis from a respected source in the AI space. I guess I'll be using slopocalypse a lot more now :) > I am bracing for 2026 as the year of the slopacolypse across all of github, substack, arxiv, X/instagram, and generally all digital media It has arrived. Github will be most affected thanks to git-terrorists at Apna College refusing to take down that stupid tutorial. IYKYK.
- ActorNightly 8mo agoThe respect is unwarranted. He ran Teslas ML division, but still doesnt know what a simple kalman filter is (in the sense where he claimed that lidar would be hard to integrate with cameras).
- akoboldfrying 8mo agoThe Kalman filter examples I've seen always involve estimating a very simple quantity, like the location of a single 3D point, from noisy sensors. It's clear how multiple estimates can be combined into a new estimate. I'd guess that cameras on a self-driving car are trying to estimate something much more complex, something like 3D surfaces labeled with categories ("person", "traffic light", etc.). It's not obvious to me how estimates of such things from multiple sensors and predictions can be sensibly and efficiently combined to produce a better estimate. For example, what if there is a near red object in front of a distant red background, so that the camera estimates just a single object, but the lidar sees two?
- ActorNightly 8mo agohttps://www.bzarg.com/p/how-a-kalman-filter-works-in-pictures/ https://www.bzarg.com/p/how-a-kalman-filter-works-in-picture... Kalman filters basic concept is essentially this. 1. make prediction on the next state change of some measurable n dimentional quantity, and estimate the covariance matrix across those n dimentions, which describe essentially a probability that the i-th dimention is going to increase (or decrease) with j-th dimention, where i and j are between 0 and n (indices of the vector) 2. Gather sensor data (that can be noisy), and reconcile the predicted measurement with the measured to get the best guess. The covariance matrix acts as a kind of weight for each of the elements 3. Update the covariance matrix based on the measurements in previous step. You can do this for any vector of numbers. For example, instead of tracking individual objects, you can have a grid where each element represents a physical object that the car should not drive into, with a value representing certainty of that object being there. Then when you combine sensor reading, you still can use your vision model but that model would be enhanced by what lidar detects, both in terms of seeing things that camera doesn't pick up and rejecting things that aren't there. And the concept is generic enough to where you can set up a system to be able to plug in any additional sensor with its own noise, and it all works out in the end. This is used all the You can even extend the concept past Gaussian noise and linearity, there are a number of other filters that deal with that, broadly under the umbrella of sensor fusion. The problem is that Karpathy is more of a computer scientist, so he is on his Code 2.0 train of having ML models do everything. I dunno if he is like that himself or Musks "im smarter than everyone else that came before me" rubbed off. And of course when you think like that, its going to be difficult to integrate lidar into the model. But the problem with that thinking is that forward inference LLM is not AI, and it will never ever be able to drive a car well compared to a true "reasoning" AI with feedback loops.
- sota_pop 8mo ago> Slopacolypse Really… REALLY not looking forward to getting this word spammed at me the next 6-12 months… even less so seeing the actual manifestation. > TLDR This should be at the start? I actually have been thinking of trying out ClaudeCode/OpenCode over this past week… can anyone provide experience, tips, tricks, ref docs? My normal workflow is using Free-tier ChatGPT to help me interrogate or plan my solution/ approach or to understand some docs/syntax/best practice of which I’m not familiar. then doing the implementation myself.
- gverrilla 8mo agoClaude code official docs are quite nice - that's where I started.
- vinhnx 8mo agoBoris Cherny (Claude Code creator) replies to Andrej Karpathy https://xcancel.com/bcherny/status/2015979257038831967 https://xcancel.com/bcherny/status/2015979257038831967
- jwilliams 8mo ago> It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. This is true... Equally I've seen it dive into a rabbit hole, make some changes that probably aren't the right direction... and then keep digging. This is way more likely with Sonnet, Opus seems to be better at avoiding it. Sonnet would happily modify every file in the codebase trying to get a type error to go away. If I prompt "wait, are you off track?" it can usually course correct. Again, Opus seems way better at that part too. Admittedly this has improved a lot lately overall.
- gregjor 8mo agoI don't understand why anyone finds it interesting that a machine, or chatbot, never tires or gets demoralized. You have to anthromorphize the LLM before you can even think of those possibilities. A tractor never tires or gets demoralized either, because it can't. Chatbots don't "dive into a rabbit hole ... and then keep digging" because they have superhuman tenacity, they do it because that's what software does. If I ask my laptop to compute the millionth Fibonacci number it doesn't sigh and complain, and I don't think it shows any special qualities unless I compare it to a person given the same job.
- akoboldfrying 8mo agoYou're a machine. You're literally a wet, analog device converting some forms of energy into other forms just like any other machine as you work, rest, type out HN comments, etc. There is nothing special about the carbon atoms in your body -- there's no metadata attached to them marking them out as belonging to a Living Person. Other living-person-machines treat "you" differently than other clusters of atoms only because evolution has taught us that doing so is a mutually beneficial social convention. So, since you're just a machine, any text you generate should be uninteresting to me -- correct? Alternatively, could it be that a sufficiently complex and intricate machine can be interesting to observe in its own right?
- felineflock 8mo agoxcancel? What is the purpose or benefit of providing a free mirror to x? Doesn't it end up sparing the x servers and causing their costs to decrease?
- energy123 8mo agoA big wow moment coming up is going to be GPT 5.* in Codex with Cerebras doing inference. The inference speed is going to be a big unlock, because many tasks are intrinsically serial. It's going to feel literally like playing God, where you type in what you want and it happens ~instantly.
- brcmthrowaway 8mo agoWhen?
- energy123 8mo agoI don't know when but I'm going off: - "OpenAI is partnering with Cerebras to add 750MW of ultra low-latency AI compute" - Sam Altman saying that users want faster inference more than lower cost in his interview. - My understanding that many tasks are serial in nature.
- cactusplant7374 8mo agoSpeed is really important to me but also I would like higher weekly limits -- which means lower cost I suppose. Building out complex projects can take 6 months to a year on a Pro plan.
- energy123 8mo agoSame experience with Pro. My trick is to attach the codebase as a txt file to 5-10 different GPT 5.2 Thinking chats, paste in the specs, and then get hard work done there, then just copy paste the final task list into codex to lower codex usage.
- erelong 8mo ago> 80% agent coding A lot of these things sound cool but sometimes I'm curious what they're actually building Like, is their bottleneck creativity now then? Are they building naything interedting or using agents to build... things that don't appeal to me, anyway?
- ewidar 8mo agoI guess it depends what appeal to you. As an example finding myself in a similar 80% situation, over the last few months I built - a personal website with my projects and poems - an app to rework recipes in a format I like from any source (text, video,...) - a 3d visual version of a project my nephew did for work - a gym class finder in my area with filters the websites don't provide - a football data game - working on a saas for work so typical saas stuff I was never that productive on personal projects, so this is great for me. Also the coding part of these projects was not very appealing to me, only the output, so it fits well with AI using. In the meanwhile I did Advent of Code as usual for the fun of code. Different objectives.
- gloosx 8mo agoSo what is he even coding there all the time? Does anybody have any info on what he is actually working on besides all the vibe-coding tweets? There seems to be zero output from they guy for the past 2 years (except tweets)
- augment_me 8mo agoThis is the first question I ask, and every time I get the answer of some monolith that supposedly solves something. Imo, this is completely fine for any personal thing, I am happy when someone says they made an API to compare weekly shopping prices from the stores around them, or some recipe, this makes sense. However more often than not, someone is just building a monolithic construction that will never be looked at again. For example, someone found that HuggingFace dataloader was slow for some type of file size in combination with some disk. What does this warrant? A 300000+ line non-reviewed repo to fix this issue. Not a 200-line PR to HuggingFace, no you need to generate 20% of the existing repo and then slap your thing on there. For me this is puzzling, because what is this for? Who is this for? Usually people built these things for practice, but now its generated, so its not for practice because you made very little effort on it. The only thing I can see that its some type of competence signaling, but here again, if the engineer/manager looking knows that this is generated, it does not have the type of value that would come with such signaling. Either I am naive and people still look at these repos and go "whoa this is amazing", or it's some kind of induced egotrip/delusion where the LLM has convinced you that you are the best builder.
- ruszki 8mo agoI don’t know, but it’s interesting that he and many others come up with this “we should act like LLMs are junior devs”. There is a reason why most junior devs work on fairly separate parts of products, most of the time parts which can be removed or replaced easily, and not an integral part of products: because their code is usually quite bad. Like every few lines contains issues, suboptimal solutions, and full with architectural problems. You basically never trust junior devs with core product features. Yet, we should pretend that an “LLM junior dev” is somehow different. These just signal to me that these people don’t work on serious code.
- 8mo ago
- jeffreygoesto 8mo ago> How much of society is bottlenecked by digital knowledge work? I think not much. The real society bottleneck is that a growing number of peeps try to convince each other that life and society are a zero sum game. They are so much more if we don't do that.
- huflungdung 8mo ago[dead]
- solarized 8mo agoNext milestone: solving authoritarian LLM dependencies. We can’t always get trapped in local minima. Or is that actually okay?
- markb139 8mo agoI retired from paid sw dev work in 2020 when COVID arrived. I’ve worked on my small projects since with all development by hand. I’d followed the rise of AI, but not used it. Late last year I started a project that included reverse engineering some firmware that runs on an Intel 8096 based embedded processor. I’d never worked on that processor before. There are tools available, but they cost many $. So, I started to think about a simple disassembler. 2 weeks ago we decided to try Claude to see what it could do. We now have a disassembler, assembler and a partially working emulator. No doubt there are bugs and missing features and the code is a bit messy, but boy has it sped up the work. One thing did occur to me. Vendors of small utilities could be in trouble. For example I needed to cut out some pages from a pdf. I could have found a tool online(I’m sure there are several), write one myself. However, Claude quickly performed the task.
- TeMPOraL 8mo ago> Vendors of small utilities could be in trouble. For example I needed to cut out some pages from a pdf. I could have found a tool online(I’m sure there are several), write one myself. However, Claude quickly performed the task. Definitely. Making small, single-purpose utilities with LLMs is almost as easy these days as googling for them on-line - much easier, in fact, if you account for time spent filtering out all the malware, adware, "to finish the process, register an account" and plain broken "tools" that dominate SERP. Case in point, last time my wife needed to generate a few QR codes for some printouts for an NGO event, I just had LLM make one as a static, single-page client-side tool and hosted it myself -- because that was the fastest way to guarantee it's fast, reliable, free of surveillance economy bullshit, and doesn't employ URL shorteners (surprisingly common pattern that sometimes becomes a nasty problem down the line; see e.g. a high-profile case of some QR codes on food products leading to porn sites after shortlink got recycled).
- Antibabelic 8mo agoWhatever happened to just typing "apt install qrencode"? It's definitely "fast, reliable, free of surveillance economy bullshit, and doesn't employ URL shorteners".
- arh5451 8mo agoThank you for the really excellent summation. I echo your thought 1 to 1. I have found it more difficult to learn new languages or coding skills, because I am no longer forced to go through the painful slow grind of learning.
- gregjor 8mo agoPainful slow grind? I have always found the learning part what I enjoy most about programming. I don't intend to outsource that a chatbot.
- ed_mercer 8mo agoDoes one ever still need to learn new languages or coding skills if an AI will be able to do it?
- dag11 8mo agoThis question makes me unbelievably sad. Why should anyone learn anything? I'm not disagreeing.
- FeteCommuniste 8mo agoProbably not. But as someone who has learned a few languages, having to outsource a conversation to a machine will never not feel incredibly lame. I doubt most people feel the same, though.
- doe88 8mo agoAre there good guides about how to write Agents or good repos with examples? Also, are there big differences between how you would write one in Codex cli vs Claude code? Can there be run on it interchangeably?
- poszlem 8mo agoI keep thinking about the TechnoCore from Dan Simmons' Hyperion, where the AIs were serving humans but secretly that was a parasitic relation, where they've been secretly using human brains as distributed processing nodes, essentially harvesting humanity's neural activity for their own computational needs without anyone's knowledge. I know this is SF, but to me working with those LLMs feels more and more like that, and the atrophy part is real. Not that the model is literally using our brains as compute, but the relationship can become lopsided.
- dubeye 8mo agothe xcancel link is amusing. 9/10 of the most important social media users use X, like or loath it
- gregorygoc 8mo agoStop whining
- dubeye 8mo agonot 100% clear who this is directed at ;)
- MORPHOICES 8mo ago[dead]
- bob1029 8mo agoI would agree that OAIs GPT-5 family of models is a phase change over GPT-4. In the ChatGPT product this is not immediately obvious and many people would strongly argue their preference for 4. However, once you introduce several complex tools and make tool calling mandatory, the difference becomes stark. I've got an agent loop that will fail nearly every time on GPT-4. It works sometimes, but definitely not enough to go to production. GPT-5 with reasoning set to minimal works 100% of the time. $200 worth of tokens and it still hasn't failed to select the proper sequence of tools. It sometimes gets the arguments to the tools incorrect, but it's always holding the right ones now. I was very skeptical based upon prior experience but flipping between the models makes it clear there has been recent stepwise progress. I'll probably be $500 deep in tokens before the end of the month. I could barely go $20 before I called bullshit on this stuff last time.
- alansaber 8mo agoPretty sure there wasn't extensive training on tooling beforehand. I mean, god, during GPT-3 even getting a reliable json output was a battle and there were dedicated packages for json inference.
- theshrike79 8mo agoNow imagine local models with 95%+ reliable tool calling, you can do insane things when that's the reality.
- TrackerFF 8mo agoMinor nitpick: The original measure of a 10x programmer was not the productivity multiplier max/mean, but rather max/min.
- jbjbjbjb 8mo ago> do generalists outperform specialists? Depends what we mean by specialist. If it frontend vs backend then maybe. If it general dev vs some specialist scientific programmer or other field where a generalist won’t have a clue then this seems like a recipe for disaster (literal disasters included).
- MarginalGainz 8mo ago[dead]
- tariky 8mo agoI used CC in year age and it was not good. But one month ago I paid for max and started to rebuild my company web shop using it. It is like plowing land with hand one year age and now is like I'm in brend new John Deere. It's amazing. Of course its not perfect but if you understand code and problem it needs to solve then it works really good.
- dzonga 8mo agomaybe its just me doing stuff that's out the usual loop even dealing with api's that have MCP servers the so called agents make a mess of everything. my stuff is just regular data stuff - ingest data from x - transform it | make it real time - then pipe it to y
- elif 8mo agoWhy am I not surprised that a blog was written about LLM coding going from 20% to 80% useful, yet all of the HN comments are still nit picking about some negative details rather than building positive ideas toward some progress... Is the programmer ego really this fragile? At least luddites had an ideological reasoning, whereas here we just seem to have emotional reflexes.
- phito 8mo agoIt's because we see a bunch of people completely ignoring the missing 20% and flooding the world with complete slop. The push back is required to keep us sane, we need people reminding others that it's not at 100% yet even if it sometimes feels like it.
- hollowturtle 8mo agoThen you have Anthropic that states on his own blog that engineers fully delegate to claude code only from 0 to 20% https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic https://www.anthropic.com/research/how-ai-is-transforming-wo... The fact that people keep pushing figures like 80% is total bs to me
- an0malous 8mo agoIt’s usually people doing side projects or non-programmers who can’t tell the code is slop. None of these vibe coding evangelists ever shares the code they’re so amazed by, even though by their own logic anyone should be able to generate the same code with AI.
- bob1029 8mo agoThis kind of thought policing is getting to be exhausting. Perhaps we need a different kind of push back. Do you know what my use case is? Do you know what kind of success rate I would actually achieve right now? Please show me where my missing 20% resides.
- 8mo ago
- jliptzin 8mo agoThe tenacity part is definitely true. I told it to keep trying when it kept getting stuck trying to spin up an Amazon Fargate service. I could feel its pain, and wanted to help, but I wanted to see whether the LLM could free itself from the thorny and treacherous AWS documentation forest. After a few dozen attempts and probably 50 KWh of energy it finally got it working, I was impressed. I could have done it faster myself, but the tradeoff would have been much higher blood pressure. Instead I relaxed and watched youtube while the LLM did its work.
- noisy_boy 8mo ago> I am bracing for 2026 as the year of the slopacolypse across all of github, substack, arxiv, X/instagram, and generally all digital media. 2026 is just when it picks up - it'll get exponentially worse. I think 2026 is the year of Business Analysts who were unable to code. Now CC et all are good enough that they can realize the vision as long as one knows exactly the requirements (software design not that important). Programmers who didn't know business could get by so far. Not anymore, because with these tools, the guy who knows business can now code fairly well.
- kitd 8mo agowith these tools, the guy who knows business can now code fairly well. ... until CC doesn't get it quite right and the guy who knows business doesn't know code.
- rubzah 8mo agoThe future of the programmer profession: This AI-generated mess of a codebase does 80% of what I want. Now fix the last 20%, should be easy, right?
- AnimalMuppet 8mo agoApart from the "AI-generated mess" part, that's too often been the past of the programmer profession, too.
- HugoDz 8mo agoAgree here, the code barrier (creating software) was hiding the real mountain: creating software business. The two are very different beasts.
- jazzpote 8mo agoSoftware engineers were never paid to "create software business". That's the job of business ppl, and a thin minority of software engineers.
- cmrdporcupine 8mo agoRight on especially on two things -- 1) the tools doing a disservice by not interviewing and seeking input and 2) The 2026 "Slopocalypse" I'm hopeful that 2026 will be the year that the biggest adopters are forced to deal with the mass of product they've created that they don't fully understand, and a push for better tooling is the result. Today's agentic tools are crude from a UX POV from where I am hoping they will end up.
- netcraft 8mo ago> Tenacity. It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. This is true to an extent for sure and they will go much longer than most engineers without getting "tired", but I've def seen both sonnet and opus give up multiple times. They've updated code to skip tests they couldn't get to pass, given up on bugs they couldn't track down, etc. I literally had it ask "could we work on something else and come back to this"
- lucianbr 8mo agoThe glorified autocomplete. Why would the LLM "work on something else then get back on this", is it's subconscious going to solve the problem during that time? But because people say it, it says it too. Making sense is optional.
- havefunbesafe 8mo agoIve found that clearing the context and getting back to it later actually DOES work. When you restart, your personal context is cleared and you might be better at describing the problem you are solving in a more informationally dense way.
- Davidzheng 8mo agonot impossible right? the new context can provide some needed hints, etc...
- Schlagbohrer 8mo agoReminiscent of a time just a year or two ago where the LLMs would get downright frustrated and sassy
- manbash 8mo agoOh, definitely. Also, they end up getting stuck in a loop, adding and removing the same code endlessly. And then someone comes and "improves" their agent with additional "do not repeat yourself" prompts scattered all over the place, to no avail. "Asinine" describes my experience perfectly.
- upghost 8mo agotl;dr - All this AI stuff is just Universal Paperclips[1] I see a lot of comments about folks being worried about going soft, getting brain rot, or losing the fun part of coding. As far as I'm concerned this is a bigger (albeit kinda flakey) self-driving tractor. Yeah I'd be bored if I just stuck to my one little cabbage patch I'd been tilling by hand. But my new cabbage patch is now a megafarm. Subjectively, same level of effort. [1]: https://en.wikipedia.org/wiki/Universal_Paperclips https://en.wikipedia.org/wiki/Universal_Paperclips
- jedisct1 8mo agoClaude is good at writing code, not so good at reasoning, and I would never trust or deploy to production something solely written by Claude. GPT-5.2 is not as good for coding, but much better at thinking and finding bugs, inconsistencies and edge cases. The only decent way I found to use AI agents is by doing multiple steps between Claude and GPT, asking GPT to review every step of every plan and every single code change from Claude, and manually reviewing and tweaking questions and responses both way, until all the parties, including myself, agree. I also sometimes introduce other models like Qwen and K2 in the mix, for a different perspective. And gosh, by doing so you immediately realize how dumb, unreliable and dangerous code generated by Claude alone is. It's a slow and expensive process and at the end of the day, it doesn't save me time at all. But, perhaps counterintuitively, it gives me more confidence in the end result. The code is guaranteed to have tons of tests and assurance for edge cases that I may not have thought about.
- rubzah 8mo agoThe Slopocalypse - an unexpected variant of Gray Goo: https://en.wikipedia.org/wiki/Gray_goo https://en.wikipedia.org/wiki/Gray_goo
- AnimalMuppet 8mo agoWell, it may consume the AI environment. Maybe even the internet. It's not going to consume a PC with g++, though (at least if the PC doesn't update g++ any more once g++ starts accepting AI contributions). There may come a point where having a "survivor machine" with auto-update turned off may be a really good idea.
- Applejinx 8mo agoI already do this, in the form of survivor machines made to do initial coding on a retro platform so the result will translate across all possible platforms. Got to, as I'm an Apple coder primarily, so if I want to target older machines I can only do it through a survivor machine: support is always pruned out of Xcode and it would be insane to try and patch it to keep everything in scope.
- direwolf20 8mo agoAslopalypse, a slop ellipse.
- svara 8mo agoBasically mirrors my experience. Interestingly, when you point out this ... > IDEs/agent swarms/fallability. Both the "no need for IDE anymore" hype and the "agent swarm" hype is imo too much for right now. The models definitely still make mistakes and if you have any code you actually care about I would watch them like a hawk, in a nice large IDE on the side. ... here on HN [0] you get a bunch of people telling you to get with the times, grandpa. Really makes me wonder: Who are these people and why are they doing that? [0] https://news.ycombinator.com/item?id=46745039 https://news.ycombinator.com/item?id=46745039
- Aperocky 8mo agoIs it really brain atrophy if I never learned to code in ASM in my entire career as compiler has been doing that for me? A part of me really want to say yes and wear it as a badge to have been coding before LLMs were a thing, but at the same time, it's not unprecedented.
- ex-aws-dude 8mo agoThe thing is the compiler does exactly what you want it to 99.999…% of the time so you never have to drop down into ASM That’s not really true in this case I think a person with zero coding knowledge would have a lot tougher time using these tools successfully
- direwolf20 8mo agoIs it muscle atrophy if you were a weakling since birth? Is it retina degeneration if you were born blind? No, because atrophy is a loss of a prior strength, and not an ever–existing weakness, but it's just as bad.
- lingrush4 8mo agoNo idea why the poster wants to deprive this author of engagement on his post, but here's the original link: https://x.com/karpathy/status/2015883857489522876 https://x.com/karpathy/status/2015883857489522876
- globular-toast 8mo ago> LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building. Who doesn't like building? Building without any thought is literally a toy, like Lego or paint by numbers. That's the entire reason those things are popular. But a game is not a job. Sometimes I feel like half the people in this career are children. Never had any real responsibility. "Oh, everyone writes bugs, who tf cares". "Move fast, break stuff" was literally and unironically the tag line for a company that should have been taking far more responsibility. This trend isn't limited to programmers either. Wherever I look I see people not taking responsibility. Lots of children in adult bodies. I do hope there are some adults who are really pulling the strings somewhere...
- 1970-01-01 8mo ago>Tenacity I've seen the exact opposite with Claude. It literally ditched my request mid-analysis when doing a root cause analysis. It decided I was tired of the service failing and then gave me some restart commands to 'just get it working'
- daxfohl 8mo agoNow that it's real, is there a minimum bar of non-AI-generated code that should be required in any production product? Like if 100% of the code is AI generated (or even doom-tabbed) and something goes wrong in prod, (crash, record corruption, data leak, whatever) then what? 99%? 50%? What's the bar where the risk starts outweighing the reward? When do we look around and say "maybe we should start slowing down before we do something that destroys our company"? Granted it's not a one-size-fits-all problem, but I'm curious if any teams have started setting up additional concrete safeguards or processes to mitigate that specific threat. It feels like a ticking time bomb. It almost begs the question, what even is the reward? A degradation of your engineering team's engineering fundamentals, in return for...are we actually shipping faster?
- cagenut 8mo agoobviously you're not a devops eng, I think you're wildly under-estimating how much of business critical code pre-ai is completely orphaned anyway. the people who wrote it were contractors long gone, or employees that have moved companies/departments/roles, or of projects that were long since wrapped up, or of people who got laid off, or the people who wrote it simply barely understood it in the first place and certainly don't remember what they were thinking back then now. basically "what moron wrote this insane mess... oh me" is the default state of production code anyway. there's really no quality bar already.
- daxfohl 8mo agoI am a devops engineer and understand your point. But there's a huge difference: legacy code doesn't change. Yeah occasionally something weird will happen and you've got to dig into it, but it's pretty rare, and usually something like an expired certificate, not a logic bug. What we're entering, if this comes to fruition, is a whole new era where massive amounts of code changes that engineers are vaguely familiar with are going to be deployed at a much faster pace than anything we've ever seen before. That's a whole different ballgame than the management of a few legacy services.
- cagenut 8mo ago
- borroka 8mo agoI am developing a web application for a dictionary that translates words from the national language into the local dialect. Vibe coding and other tools, such as Google Vision, helped me download images published online, compile a PDF, perform OCR (Tesseract and Google Vision), and save everything in text format. The OCR process was satisfactory for a first draft, but the text file has a lot of errors, as you'd expect when the dictionary has about 30,000 entries: Diacritical marks disappear, along with typographical marks and dashes, lines are moved up and down, and parts of speech (POS) are written in so many different ways due to errors that it is necessary to identify the wrong POS's one by one. If the reasoning abilities of LLM-derived coding agents were as advanced as some claim, it would be possible for the LLM to derive the rules that must be applied to the entire dictionary from a sufficiently large set of “gold standard” examples. If only that were the case. Every general rule applied creates other errors that propagate throughout the text, so that for every problem partially solved, two more emerge. What is evident to me is not clear to the LLM, in the sense that it is simple for me, albeit long and tedious, to do the editing work manually. To give an example, if trans.v. (for example) indicates a transitive verb, it is clear to me that .trans.v. is a typographical error. I can tell the coding tool (I used Gemini, Claude, and Codex, with Codex being the best) that, given a standard POS, if there is a “.” before it, it must be deleted because it is a typo. The generalization that comes easily to me but not to the coding agent is that if not one but two periods precede the POS, it means there are two typos, not to delete just one of the two dots. This means that almost all rules have to be specified, whereas I expected the coding agent to generalize from the gigantic corpus on which it was trained (it should “understand” what the POS are, typical typos, the language in which the dictionary is written, etc.). The transition from text to json to webapp is almost miraculous, but what is still missing from the mix is human-level reasoning and common sense (in part, I still believe that coding agents are fantastic, to be clear).
- jermberj 8mo ago> The most common category is that the models make wrong assumptions on your behalf and just run along with them without checking. They also don't manage their confusion, they don't seek clarifications, they don't surface inconsistencies, they don't present tradeoffs, they don't push back when they should, and they are still a little too sycophantic. Does this not undercut everything going on here. Like, what?
- awsanswers 8mo agoIt's predictable so you run defense around it with prompting, validation and model tuning. It generates volumes of working code in seconds from natural language prompts so it's extremely business efficient. We're talking about tools that generate correct code to 95% of a solution, the follow up human and automated test review, and second coding pass to fix the 5% are a non issue.
- enduser 8mo agoThe way I have managed junior engineers is 90% via PR and testing, 10% via reading code in an editor or IDE. It’s hard to let go of being the keyboard jockey, but in so many cases it is better to describe plans and acceptance criteria and just review the diffs.
- Jean-Papoulos 8mo ago>Generation (writing code) and discrimination (reading code) are different capabilities in the brain. Largely due to all the little mostly syntactic details involved in programming, you can review code just fine even if you struggle to write it. If this is how all juniors are learning nowadays, seniors are going shot up in value in the next decade.
- seffietron 8mo agoAt the risk of exposing my... atypical take on contemporary occupational "morality". LLMs, and Claude Code specifically, have given me the ability to work two jobs, and retain my soft-standing as the "guy" that gets stuff done/knows my stuff/can fix anything, while still working less hours than I did when I just had a single job. I firmly believe, at SOME point, ML is going to eat my lunch. And I'd like to be well and retired off to a countryside homestead by then. Until such a time, I am going to use and abuse this technology as much as possible to gather 4 paychecks a month, optimize my investment portfolio to scale my NW, and by any means gain financial independence before the risk of my career vaporizing materializes. Sure, I could REALLY try and be one of those engineers that pulls a $750k salary and I wouldn't need to do this; but that isn't really in the cards for me. I know where I stand and I'm simply not smart or hardworking enough to get paid that much from a single job and guarantee my financial independence in the traditional way. To that end, these tools have been extraordinarily impactful for me in a very simple and objectively positive way. And as much as I relate to basically everything OP mentioned, at the end of the day I simply DGAF. I want to make as much money as possible before the music stops, and this is the smart way to do that right now
- rikdom 8mo agoHonestly man, this is totally understandable. It's a rat race, and if you don't use the tools at your disposal, you'll be left behind.
- trivo 8mo agoI sometimes wonder about the similarities between this paradigm switch (coding -> vibe coding) and when the industry switched from writing assembler to using high-level languages. I both cases we switched from having to specify every posibble implementation detail to focusing more on higher level concepts and letting the machine work out the rest. Maybe in the future instead of sharing source code, we will share prompts that we used to create a program. Similarly how different compilers produce different assembly now, "compiling" prompts with different agent/model would give different results. Maybe in the future an analog for "optimizing compiler" would emerge for agents, which would turn the (working) slop into something more clean.