8 ms·
A year of vibes
- bgwalter 10mo agoIt is nice that he speaks about some of the downsides as well. In many respects 2025 was a lost year for programming. People speak about tools, setups and prompts instead of algorithms, applications and architecture. People who are not convinced are forced to speak against the new bureaucratic madness in the same way that they are forced to speak against EU ChatControl. I think 2025 was less productive, certainly for open source, except that enthusiasts now pay the Anthropic tax (to use the term that was previously used for Windows being preinstalled on machines).
- grim_io 10mo agoI'm glad there has been a break in endless bikeshedding over TDD, OOP, ORM(partially) and similar.
- r2_pilot 10mo ago>>"I think 2025 was less productive" I think 2025 is more productive for me based on measurable metrics such as code contribution to my projects, better ability to ingest and act upon information, and generally I appreciate the Anthropic tax because Claude genuinely has been a step-change improvement in my life.
- moooo99 10mo ago> more productive for me based on measurable metrics such as code contribution to my projects Isn‘t it generally agreed upon that counting contributions, LoC or similar metrics is a very bad way to gauge productivity?
- r2_pilot 10mo agoI don't care about industry metrics when I'm building my own AI research robotics platform and it's doing what I ask it to do, proving itself in the real world far better than any performative best-practice theatrics in the service of risible MBA-grade effluvia masquerading as critical discourse.
- theshrike79 10mo agoDuring 2025 I've almost exhausted my personal TODO-list of small applications and created a few extra ones. This would've never happened without a Claude Pro (+ChatGPT) subscription. And as I'm not American, none of them are aimed to be subscription based SaaS offerings, they're just simple CLI applications for my personal use. If someone else enjoys them, good for them. =)
- JimDabell 10mo ago> In many respects 2025 was a lost year for programming. People speak about tools, setups and prompts instead of algorithms, applications and architecture. I think the opposite. Natural language is the most significant new programming language in years, and this year has had a tremendous amount of progress in collectively figuring out how to use this new programming language effectively.
- 9rx 10mo ago> and this year has had a tremendous amount of progress in collectively figuring out how to use this new programming language effectively. Hence the lost year. Instead of productively building things, we spent a lot of resources on trying to figure out how to build things.
- sixtyj 10mo agoAbsolutely. So much noise. "There’s an AI for that" lists 44,172 AI tools for 11,349 tasks. Most of them are probably just wrappers… As Cory Doctorow uses enshittification for the internet, for AI/LLM there should be something like a dumbaification. It reminds me late 90s when everything was "World Wide Web". :) Gold rush it is.
- theshrike79 10mo agoThis is just the Blockchain and Web3 (NTF) crazes all over again on the surface. Every single grifter from those times is slapping AI on everything that moves or doesn't move. But the difference is that the blockchain was (and still is) a solution looking for a problem. LLMs can solve actual problems today.
- wiseowise 10mo ago> algorithms, applications and architecture. Which one is that? Endless leetcode madness? Or constant bikeshedding about today's flavor of MVC (MVI, MVVM, MVVMI) or whatever else bullshit people come up with instead of actually shipping?
- data-ottawa 10mo agoMaybe it's because I'm a data scientist and not a dedicated programmer/engineer, but setup+tooling gains this year have made 2025 a stellar year for me. DS tooling feels like it hit much a needed 2.0 this year. Tools are faster, easier, more reliable, and more reproducible. Polars+pyarrow+ibis have replaced most of my pandas usage. UDFs were the thing holding me back from these tools, this year polars hit the sweet spot there and it's been awesome to work with. Marimo has made notebooks into apps. They're easier to deploy, and I can use anywidget+llms to build super interactive visualizations. I build a lot of internal tools on this stack now and it actually just works. PyMC uses jax under the hood now, so my MCMC workflows are GPU accelerated. All this tooling improvement means I can do more, faster, cheaper, and with higher quality. I should probably write a blog post on this.
- JKCalhoun 10mo agoGot distracted: love the "WebGL metaballs" header and footer on the site.
- simonw 10mo agoI really feel this bit: > With agentic coding, part of what makes the models work today is knowing the mistakes. If you steer it back to an earlier state, you want the tool to remember what went wrong. There is, for lack of a better word, value in failures. As humans we might also benefit from knowing the paths that did not lead us anywhere, but for machines this is critical information. You notice this when you are trying to compress the conversation history. Discarding the paths that led you astray means that the model will try the same mistakes again. I've been trying to find the best ways to record and publish my coding agent sessions so I can link to them in commit messages, because increasingly the work I do IS those agent sessions. Claude Code defaults to expiring those records after 30 days! Here's how to turn that off: https://simonwillison.net/2025/Oct/22/claude-code-logs/ https://simonwillison.net/2025/Oct/22/claude-code-logs/ I share most of my coding agent sessions through copying and pasting my terminal session like this: https://gistpreview.github.io/?9b48fd3f8b99a204ba2180af785c89d2 https://gistpreview.github.io/?9b48fd3f8b99a204ba2180af785c8... - via this tool: https://simonwillison.net/2025/Oct/23/claude-code-for-web-video/#the-result https://simonwillison.net/2025/Oct/23/claude-code-for-web-vi... Recently been building new timeline sharing tools that render the session logs directly - here's my Codex CLI one (showing the transcript from when I built it): https://tools.simonwillison.net/codex-timeline?url=https%3A%2F%2Fgist.githubusercontent.com%2Fsimonw%2Fb9bc2cc5b5d996de3f7383796180276c%2Fraw%2Fb9af382e14bf5ab19aad3d23d4d788dbeef54aa1%2Ftimeline.jsonl#tz=local&q=&type=all&payload=all&role=all&hide=1&truncate=1 https://tools.simonwillison.net/codex-timeline?url=https%3A%... And my similar tool for Claude Code: https://tools.simonwillison.net/claude-code-timeline?url=https://gist.githubusercontent.com/simonw/3b2dc0fbea4ba8e677bf6195bc8d9952/raw/85d69ad9f1b874c2139a2f61c7b0f18ed30856c4/4e485649-3221-4d7b-bf1a-129dff975689.jsonl#tz=local&q=&type=all&ct=all&role=all&hide=0&truncate=1&sel=3 https://tools.simonwillison.net/claude-code-timeline?url=htt... What I really want it first class support for this from the coding agent tools themselves. Give me a "share a link to this session" button!
- stacktraceyo 10mo agoI’d like to make something like this but in the background. So I can better search my history of sessions. Basically start creating my own knowledge base of sorts
- divbzero 10mo ago> My biggest unexpected finding: we’re hitting limits of traditional tools for sharing code. The pull request model on GitHub doesn’t carry enough information to review AI generated code properly — I wish I could see the prompts that led to changes. It’s not just GitHub, it’s also git that is lacking. The limits seem to be not just in the pull request model on GitHub, but also the conventions around how often and what context gets committed to Git by AI. We already have AGENTS.md (or CLAUDE.md, GEMINI.md, .github/copilot-instructions.md) for repository-level context. More frequent commits and commit-level context could aid in reviewing AI generated code properly.
- tolerance 10mo agoArmin has some interesting thoughts about the current social climate. There was a point where I even considered sending a cold e-mail and asking him to write more about them. So I’m looking forward to his writing for Dark Thoughts—the separate blog he mentions.
- kashyapc 10mo ago"Because LLMs now not only help me program, I'm starting to rethink my relationship to those machines. I increasingly find it harder not to create parasocial bonds with some of the tools I use. I find this odd and discomforting [...] I have tried to train myself for two years, to think of these models as mere token tumblers, but that reductive view does not work for me any longer." It's wild to read this bit. Of course, if it quacks like a human, it's hard to resist not quacking back. As the article says, being less reckless with the vocabulary ("agents", "general intelligence", etc) could be one way to to mitigate this. I appreciate the frank admission that the author struggled for two years. Maybe the balance of spending time with machines vs. fellow primates is out of whack. It feels dystopic to see very smart people being insidiously driven to sleep-walk into "parasocial bonds" with large language models! It reminds me of the movie Her[1], where the guy falls "madly in love with his laptop" (as the lead character's ex-wife expresses in anguish). The film was way ahead of its time. [1] https://www.imdb.com/title/tt1798709/ https://www.imdb.com/title/tt1798709/
- mlinhares 10mo agoSame here, I'm seeing more and more people getting into these interactions and wondering how long until we have widespread social issues due to these relationships like people have with "influencers" on social networks today. It feels like this situation is much more worrisome as you can actually talk to the thing and it responds to you alone, so it definitely feels like there's something there.
- mjr00 10mo agoIt helps a lot if you treat LLMs like a computer program instead of a human. It always confuses me when I see shared chats with prompts and interactions that have proper capitalization, punctuation, grammar, etc. I've never had issues getting results I've wanted with much simpler prompts like (looking at my own history here) "python grpc oneof pick field", "mysql group by mmyy of datetime", "python isinstance literal". Basically the same way I would use Google; after all, you just type in "toledo forecast" instead of "What is the weather forecast for the next week in Toledo, Ohio?", don't you? There's a lot of black magic and voodoo and assumptions that speaking in proper English with a lot of detailed language helps, and maybe it does with some models, but I suspect most of it is a result of (sub)consciously anthropomorphizing the LLM.
- mritchie712 10mo agotacking on to the "New Kind Of" section: New Kind of QA: One bottle neck I have (as a founder of a b2b saas) is testing changes. We have unit tests, we review PRs, etc. but those don't account for taste. I need to know if the feature feels right to the end user. One example: we recently changed something about our onboarding flow. I needed to create a fresh team and go thru the onboarding flow dozens of times. It involves adding third party integrations (e.g. Postgres, a CRM, etc.) and each one can behave a little different. The full process can take 5 to 10 minutes. I want an agent go thru the flow hundreds of times, trying different things (i.e. trying to break it) before I do it myself. There are some obvious things I catch on the first pass that an agent should easily identify and figure out solutions to. New Kind of "Note to Self": Many of the voice memos, Loom videos, or notes I make (and later email to myself) are feature ideas. These could be 10x better with agents. If there were a local app recording my screen while I talk thru a problem or feature, agents could be picking up all sorts of context that would improve the final note. Example: You're recording your screen and say "this drop down menu should have an option to drop the cache". An agent could be listening in, capture a screenshot of the menu, find the frontend files / functions related to caching, and trace to the backend endpoints. That single sentence would become a full spec for how to implement the feature.
- rootnod3 10mo agoSorry, but why would including the prompt in the pull request make any difference? Explain what you DID in the pull request. If you can't summarize it yourself, it means you didn't review it yourself, so why should I have to do it for you?
- theshrike79 10mo agoYou're making assumptions, of course you add BOTH. The point of adding the "prompt", or the discussion with the LLM is learning. You can go back and see what was the exact conversation.
- rootnod3 10mo agoSounds more like just adding a ton of wasted time for the reviewer to read through those discussions. At least summarize it yourself, e.g. "After discovering manpage XYZ, it became clear that the correct usage of this function is fooBar()".
- theshrike79 10mo agoWhy would the reviewer need to read through discussions? The description + code should be just fine. It's like having someone watch a livestream screen recording of you writing the code. It's nice to have there IF you need to go back and learn something, but hardly a review requirement.
- rootnod3 10mo ago"I have seen some people be quite successful with this." Wait until those people hit a snafu and have to debug something in prod after they mindlessly handed their brains and critical thinking to a water-wasting behemoth and atrophied their minds. EDIT: typo, and yes I see the irony :D
- wiseowise 10mo ago> Wait until those people hit a snafu and have to debug something in prod after they mindlessly handed their brains and critical thinking to a water-wasting behemoth and atrophied their minds. You've just described typical run of the mill company that has software. LLMs will make it easier to shoot yourself in the foot, but let's not rewrite history as if stackoverflow coders are not a thing.
- rootnod3 10mo agoDifference: companies are not pushing their employees to use stack overflow. Stack overflow doesn't waste massive amounts of water and energy. Stack overflow does not easily abuse millions of copyrights in a second by scraping without permission.
- rootnod3 10mo agoAnother difference: stack overflow tells you you are wrong or tells you and do your own research or to read the manual (which in a high percentage of cases is the right answer). It doesn't tell you that you are right and proceeds to hallucinate some non-existent flags for some command invocation.
- christophilus 10mo agoIt mostly incorrectly flags your question as a dup.
- pigpop 10mo agoThis is a problem but it's a known one which both Google and Anthropic seem to be making progress towards solving. I've had a full on argument with Gemini 3 where it turned out I was wrong and it correctly stuck to its guns and wouldn't let me convince it otherwise. It eventually got through to me about the mistake I made and I learned something useful from it. Sonnet and Opus are still a bit too happy to tell you "you're absolutely right" but I've noticed more pushback creeping in in the right places. It's a tough balance to get right, nobody wants to pay for a service that just tells them "no" whenever they want to try something silly or unconventional.
- Rakshath_1 10mo agoThe part that resonated most for me is the mismatch between agentic coding and our existing social/technical contracts (git, PRs, reviews). We’re generating more code than ever but losing visibility into how it came to be prompts, failures, local agent reviews. That missing context feels like the real bottleneck now not model quality.
- CuriouslyC 10mo agoI understand the parasocial bit. I actively dislike the idea of gooning, ERP and AI therapists/companions, but I still notice I'm lonelier and more distant on the days when I'm mostly writing/editing content rather than chatting with my agents to build something. It feels enough like interacting with a human to keep me grounded in a strange way.
- yawnr 10mo agoYou guys need to touch grass. Go join a kickball league or something.
- CuriouslyC 10mo agoI'd argue that doing something you don't like with people you're not into is a L. Loneliness isn't optimal but for some people it's the lesser evil. I'm married though, so I have a floor, I'm sure some people are lonely enough to benefit from being around people even under the worst of circumstances.
- yawnr 10mo agoSure, I’m not saying it needs to be kickball. I’m just saying if you find yourself being grounded by an LLM, maybe you should seek out a community of people you do actually like who do something you actually like.
- CuriouslyC 10mo agoI appreciate the sentiment, and I'm sure you mean well. It does feel a bit patronizing though, please consider that there are a lot of competent people who've experienced loneliness, chewed through their local meetups and Facebook events, found them wanting, and decided a little loneliness was a better choice.
- deleted 10mo ago[deleted]
- anshulbhide 10mo ago> The pull request model on GitHub doesn’t carry enough information to review AI generated code properly — I wish I could see the prompts that led to changes. It’s not just GitHub, it’s also git that is lacking. Yes! Who is building this?
- amarant 10mo agoCreate a folder called "prompts". Create a new file for each prompt you make, name the time after timestamp. Or just append to prompts.txt Either way, git will make it trivial to see which prompt belongs with which commit: it'll be in the same diff! You can write a pre-commit hook to always include the prompts in every commit, but I have a feeling most Vibe coders always commit with -a anyway
- NitpickLawyer 10mo agoIt's not just that. There's a lot of (maybe useful) info that's lost without the entire session. And even if you include a jsonl of the entire session, just seeing that is not enough. It would be nice to be able to "click" at some point and add notes / edit / re-run from there w/ changes, etc. Basically we're at a point where the agents kinda caught up to our tooling, and we need better / different UX or paradigms of sharing sessions (including context, choices, etc)
- ojr 10mo agoIn the next year developers need to realize normal people do not care about the tech stack or the tools used, there are far too many written thoughts and opinions and not enough polished deployed projects. From an industry standpoint it’s business as usual, acquihires from products that LLMs apparently couldn’t save.
- teaearlgraycold 10mo agoThey care - but only for how the tech stack affects the product quality. Show someone a bloated React site on 3G and compare their experience to an SSR competitor.
- ojr 10mo ago99% of the US population have access to 4G, caring about 3G is a wasted effort. The point I’m trying to make about tech stack is users don’t care if you used Gemini, ChatGPT or Claude to generate code. As someone who hasn’t converted to SSR yet. My main reason why I am switching is SEO, the performance increase is a plus though.
- summarity 10mo agoHere’s something else that just started to rally work this year with Opus 4.5: interacting with Ghidra. Nearly every binary is now suddenly transparent, in many cases it can navigate binaries better than source code itself. There’s even a research team that has bee using this approach to generate compilable C++ from binaries and run static analysis on it, to find more vulnerabilities than source analysis without involving dynamic tracing.
- theptip 10mo agoA really interesting point that keeps coming up in discussions about LLMs is “what trade-offs need to be re-evaluated” > I also believe that observability is up for grabs again. We now have both the need and opportunity to take advantage of it on a whole new level. Most people were not in a position where they could build their own eBPF programs, but LLMs can One of my big predictions for ‘26 is the industry following through with this line of reasoning. It’s now possible to quickly code up OSS projects of much higher utility and depth. LLMs are already great at Unix tools; a small api and codebase that does something interesting. I think we’ll see an explosion of small tools (and Skills wrapping their use) for more sophisticated roles like DevOps, and meta-Skills for how to build your own skill bundles for your internal systems and architecture. And perhaps more ambitiously, I think services like Datadog will need to change their APIs or risk being disrupted; in the short term nobody is going to be able to move fast enough inside a walled garden to keep up with the velocity the Claude + Unix tools will provide. UI tooling is nice, but it’s not optimized for agents.
- shimman 10mo agoDo you have any example repos of these OSS projects? I'm being reminded of this post every time people keep extolling how "productive" LLMs are: https://mikelovesrobots.substack.com/p/wheres-the-shovelware-why-ai-coding https://mikelovesrobots.substack.com/p/wheres-the-shovelware... Where is the resulting software?
- adamisom 10mo ago>Where is the resulting software? Everywhere. Remember Satya Nadella estimating 30% of code at Microsoft was written by AI? That was March. At this point it's ubiquitous—and invisible.
- adamisom 10mo agoThe very first thing I did vibe-coding was commit my prompts and AI responses. In Cursor that's extremely easy—just 'export' a chat. I stopped for security concerns but perhaps something like that is the way.
- zombiemama 10mo agoSecondly, if his creations are going to be relied upon, it will be the programmer's primary task to design his artifacts so understandable, that he can take the responsibility for them, and, regardless of the answer to the question how much of his current activity may ultimately be delegated to machines, we should always remember that neither "understanding" nor "being responsible" can properly be classified as activities: they are more like "states of mind" and are intrinsically incapable of being delegated. EWD 540 - https://www.cs.utexas.edu/~EWD/transcriptions/EWD05xx/EWD540.html https://www.cs.utexas.edu/~EWD/transcriptions/EWD05xx/EWD540...
- zkmon 10mo agoI spoke to a few people outside of IT and Tech recently. They are senior people running large departments at their companies. To my surprise, they do not think AI agents are going to have any impact in their businesses. The only solid use case they have for AI is a chat interface which, they think, can be very useful as an assistant helping with text and reports. So, I guss it's just us who are in the techie pit and think that everyone else is also is in the pit and use agents etc.
- theshrike79 10mo agoI think it's because what tech people do is objectively verifiable. Did the thing the agent made do what it was supposed to do? Yes/No. There's no "mayyyybe" or feelings or opinions. If the sort algorithm doesn't sort it doesn't work. But a secretary-agent for a non-techie is more about The Feels. It can summarize emails, "punch up" writing etc. But you can't measure whatever it outputs by anything except feels and opinions.
- petcat 10mo agoI respect Armin's opinions on the state-of-the-art in programming a lot. I'm wondering if he finds that "vibe coding" (or vibe engineering) is particularly pleasant and effective in Rust compared to, say, Python.
- johnwheeler 10mo agoI bet it would be probably even nicer. I've been programming DSP in C++ with JUCE. I have a very rusty C++ experience from years ago, but it's getting me through a lot of it, and I feel pretty comfortable. Maybe my ignorance is bliss, and I'm really just putting out bad shit.
- bigfishrunning 10mo ago> My biggest unexpected finding: we’re hitting limits of traditional tools for sharing code. The pull request model on GitHub doesn’t carry enough information to review AI generated code properly — I wish I could see the prompts that led to changes. It’s not just GitHub, it’s also git that is lacking. I find when submitting a complex PR, i tend to do a self review, adding another layer of comments above those that are included in the code. Seems like a nice place to stuff prompts
- netdevphoenix 10mo agoIn case people don't know, the author of the post is the creator of Flask, possibly Python's most popular micro web framework.