16 ms·
Superpowers: How I'm using coding agents in October 2025
- gjm11 1y agoHas anyone ever seen an instance in which the automated "How" removal actually improves an article title on HN rather than just making them wrong? (There probably are some. Most likely I notice the bad ones more than the good ones. But it does seem like I notice a lot of bad ones, and never any good ones.) [EDITED to add:] For context, the actual article title begins "Superpowers: How I'm using ..." and it has been auto-rewritten to "Superpowers: I'm using ...", which completely changes what "Superpowers" is understood as applying to. (The actual intention: superpowers for LLM coding agents. The meaning after the change: LLM coding agents as superpowers for humans.)
- add-sub-mul-div 1y agoI agree, I'm sure I've seen instances of where it's worked but the problem is that when it messes it up it's much more annoying than any benefit it brings when it does work. Some of us don't want to be reminded that tech is full of hubris, overconfidence, poor judgment, and failure about what can/should be abstracted and automated.
- dvfjsdhgfv 1y agoYeah, to the point I can recall several examples where the title stuck out as dumb on HN and only when visiting the original page it started to make sense, but not a single case where I could say the automated removal really did a good job.
- bryanrasmussen 1y agoI've had it happen with me a few times where it was reasonable, sometimes where it was debatable, and if it was just wrong I edit it to add the How back in.
- jvanderbot 1y agoThis is so interesting but it reads like satire. I'm sure folks who love persuading and teaching and marshalling groups are going to do very well in SWEng. According to this, we'll all be reading the feelings journals of our LLM children and scolding them for cheating on our carefully crafted exams instead of, you know, making things. We'll read psychology books, apparently. I like reading and tinkering directly. If this is real, the field is going to leave that behind.
- sunir 1y agoWe certainly will; they can’t replace humans in most language tasks without having a human like emotional model. I have a whole therapy set of agents to debug neurotic long lived agents with memory.
- jvanderbot 1y agoOk, call me crazy, but I don't actually think there's any technical reason that a theoretical code generation robot needs emotions that are as fickle and difficult to manage as humans. It's just that we designed this iteration of technology foundationally on people's fickle and emotional reddit posts among other things. It's a designed-in limitation, and kind of a happy accident it's capable of writing code at all. And clearly carries forward a lot of baggage...
- sunir 1y agoMaybe. I use QWAN frequently when working with the coding agents. That requires an llm equivalent of interoception to recognize when the model understanding is scrambled or “aligned with itself” which is what qwan is.
- ambicapter 1y agoIf you can find enough training data that does human-like things without have human-like qualities, we are all ears.
- jvanderbot 1y agoIt can be simultaneously the best we have, and well short of the best we want. It can be a remarkable achievement and fall short of the perceived goals. That's fine. Perhaps we can RL away some of this or perhaps there's something else we need. Idk, but this is the problem when engineers are the customer, designer, and target audience.
- sunir 1y agoQuality Spock pun.
- lerp-io 1y agotake #73895 on how to fix ur prompt to make ur slop better.
- anuramat 1y agois better slop a bad thing somehow?
- dvfjsdhgfv 1y agoWell, slop is slop, we can discuss the details but the basic thing is invariant.
- anuramat 1y agowhy reiterate the invariant?
- tonyedgecombe 1y agoYou can only regurgitate the meal so many times.
- apwell23 1y agoyeah none them can actually prove or even explain it in words why thier own golden prompting technique is superior. its all vibes. so annoying, i want to slap these ppl lol.
- lerp-io 1y agofor real lmao
- amelius 1y agoIt's not a superpower if everybody has that same power.
- cantor_S_drug 1y agoEveryone is better off with mobile phones. We can solve more diverse problems faster. Similarly we can combine our diverse superpowers (as they show in kids cartoons)
- Avicebron 1y agoI often feel these types of blogposts would be more helpful if they demonstrated someone using the tools to build something non-trivial. Is Claude really "learning new skills" when you feed it a book, or does it present it like that because you're prompting encourages that sort of response-behavior. I feel like it has to demo Claude with the new skills and Claude without. Maybe I'm a curmudgeon but most of these types of blogs feel like marketing pieces with the important bit is that so much is left unsaid and not shown, that it comes off like a kid trying to hype up their own work without the benefit of nuance or depth.
- khaledh 1y agoAgreed. The methodology needed here is something like an A/B test, with quantifiable metrics that demonstrate the effectiveness of the tool. And to do it not just once, but many times under different scenarios so that it demonstrates statistical significance. The most challenging part when working with coding agents is that they seem to do well initially on a small code base with low complexity. Once the codebase gets bigger with lots of non-trivial connections and patterns, they almost always experience tunnel vision when asked to do anything non-trivial, leading to increased tech debt.
- mwigdahl 1y agoThe problem is that you're talking about a multistep process where each step beyond the first depends on the particular path the agent starts down, along with human input that's going to vary at each step. I made a crude first stab at an approach that at least uses similar steps and structure to compare the effectiveness of AI agents. My approach was used on a small toy problem, but one that was complex enough the agents couldn't one-shot and required error correction. It was enough to show significant differences, but scaling this to larger projects and multiple runs would be pretty difficult. https://mattwigdahl.substack.com/p/claude-code-vs-codex-cli-head-to https://mattwigdahl.substack.com/p/claude-code-vs-codex-cli-...
- potatolicious 1y agoWhat you're getting at is the heart of the problem with the LLM hype train though, isn't it? "We should have rigorous evaluations of whether or not [thing] works." seems like an incredibly obvious thought. But in the realm of LLM-enabled use cases they're also expensive. You'd need to recruit dozens, perhaps even hundreds of developers to do this, with extensive observation and rating of the results. So rather than actually try to measure the efficacy, we just get blog posts with cherry-picked example of "LLM does something cool". Everything is just anecdata. This is also the biggest barrier to actual LLM adoption for many, many applications. The gap between "it does something REALLY IMPRESSIVE 40% of the time and shits the bed otherwise" and "production system" is a yawning chasm.
- jackblemming 1y agoSeems cute, but ultimately not very valuable without benchmarks or some kind of evaluation. For all I know, this could make Claude worse.
- jelling 1y agoSame. We've all fooled ourselves into believing that an LLM / stochastic process was finally solved based on a good result. But the sample size is always to low to be meaningful.
- anuramat 1y agoeven if it works as described, I'm assuming it's extremely model dependent (eg book prerequisites), so you'd have to re-run this for every model you use, this is basically poor man's finetuning; maybe explicit support from providers would make it feasible?
- tobbe2064 1y agoWhat's the cost of running with agents like this?
- dbbk 1y agoClaude Max is fixed cost
- tobbe2064 1y agoIs it doable with just pro?
- redhale 1y agoMy guess is: absolutely not, at least not for more than a few minutes. Subagents chew through tokens at a very high rate, and this system makes heavy use of subagents.
- fnicfnac 1y ago"20X the usage of pro" still sounds like quotas where the hammer could fall as it becomes less of an experiment for a limited number of power users.. The costs of self hosting some reasonable size models for a development group of various sizes is what I would want to know before investing in the skills to do a high usage style that might be being mostly bankrolled by investors for now.
- jmull 1y ago> <EXTREMELY_IMPORTANT>…*RIGHT NOW, go read… I don’t like the looks of that. If I used this, how soon before those instructions would be in conflict with my actual priorities? Not everything can be the first law.
- deleted 1y ago[deleted]
- apwell23 1y agodon't llm tell you not to give them instructions like that these days
- therealdrag0 1y agoSeems like maintaining a bashrc file. Sometimes you have to go tweak it.
- simonw 1y agoI can't recommend this post strongly enough. The way Jesse is using these tools is wildly more ambitious than most other people. Spend some time digging around in his https://github.com/obra/Superpowers https://github.com/obra/Superpowers repo. I wrote some notes on this last night: https://simonwillison.net/2025/Oct/10/superpowers/ https://simonwillison.net/2025/Oct/10/superpowers/
- csar 1y agoI’m curious how you think this compares to the Research -> Plan -> Implement method and prompts from the “Advanced Context Engineering from Agents” video when it comes to actual coding performance on large codebases. I think picking up skills is useful for broadening agents abilities, but I’m not sure I’d that’s the right thing for actual development. The packaged collection is very cool and so is the idea of automatically adding new abilities, but I’m not fully convinced that this concept of skills is that much better than having custom commands+sub-agents. I’ll have to play around with it these next few days and compare.
- ehsanu1 1y agoUsing Research->Plan->Implement flow is orthogonal, though I notice parts of those do exist as skills too. But you sometimes need to do other things too, e.g. debugging in the course of implementing or specific techniquws to improve brainstorming/researching. Some of these skills are probably better as programmed workflows that the LLM is forced to go through to improve reliability/consistency, that's what I've found in my own agents, rather than using English to guide the LLM and trusting it to follow the prescribed set of steps needed. Some mix of LLMs (choosing skills, executing the fuzzy parts of them) and just plain code (orchestration of skills) seems like the best bet to me and what I'm pursuing.
- drivebyhooting 1y agoOrthogonal means there should not be any overlap.
- spprashant 1y agoI am not ashamed to admit this whole agentic coding movement has moved beyond me. Not only do I have know everything about the code, data and domain, but now I need to understand this whole AI system which is a meta skill of its own. I fear I may never be able catch up till someone comes along and simplifies it for pleb consumption.
- gdulli 1y agoIt's also possible to put in enough hours of real coding to get to the point where coding really isn't that hard anymore, at least not hard enough to justify switching from those stable/solid fundamental skills to a constantly revolving ecosystem of ephemeral tools, models, model versions, best practices, lessons from trial and error, etc. Then you could bypass all of this distraction. Admittedly that stance is easiest to take if you were old enough, experienced enough already by the time this era hit.
- paweladamczuk 1y ago"There exist developers whose performance cannot be boosted by an LLM" is a really strong statement.
- gdulli 1y agoThe point is that it takes significant time and attention to keep up with the treadmill of constantly learning the new tool/model/framework of the month, so there's significant opportunity cost. I have continued putting 100% of my attention on the direct problems I'm solving. I don't see the coding as the hard or critical part of my work, so I don't put effort into accelerating or delegating that part.
- _se 1y agoNot really. It's on the people asserting the positive (that LLMs do improve productivity for sufficiently experienced engineers) to prove it. In the absence of proof, the null hypothesis is the default.
- 1y ago
- daemontus 1y agoMaybe this is a naive question, but how are "skills" different from just adding a bunch od examples of good/bad behavior into the prompt? As far as I can tell, each skill file is a bunch of good/bad examples of something. Is the difference that the model chooses when to load a certain skill into context?
- nrjames 1y agoI think it just gives you the ability to easily do that with slash command, like using "/brainstorm database schema" or something instead of needing to define what "brainstorm" means each time you want to do it.
- hackernewds 1y agowhat you are suggesting is 1-shot, 2-shot, 5-shot etc prompting which is so effective that it's how benchmarks were presented for a while
- simonw 1y agoI think that's one of the key things: skills don't take up any of the model context until the model actively seeks out and uses them. Jesse on Bluesky: https://bsky.app/profile/s.ly/post/3m2srmkergc2p https://bsky.app/profile/s.ly/post/3m2srmkergc2p > The core of it is VERY token light. It pulls in one doc of fewer than 2k tokens. As it needs bits of the process, it runs a shell script to search for them. The long end to end chat for the planning and implementation process for that todo list app was 100k tokens. > It uses subagents to manage token-heavy stuff, including all the actual implementation.
- deleted 1y ago[deleted]
- tcdent 1y agoThis style of prompting, where you set up a dire scenario in order to try to evoke some "emotional" response from the agent, is already dated. At some point, putting words like IMPORTANT in all uppercase had some measurable impact, but at the present time, models just follow instructions. Save yourself the experience of having to write and maintain prompts like this.
- bcoates 1y agoAlso the persuasion paper he links isn't at all about what he's talking about. That paper is about using persuasion prompts to overcome trained in "safety" refusals, not to improve prompt conformance.
- danshapiro 1y agoCo-Author of the paper here. We don't know exactly why modern llms don't want to call you a jerk, or for that matter why persuasive techniques convince them otherwise. it's not a hard line like many of the guardrails. That said, I talked to Jesse about this, and I strongly suspect the same techniques will work for prompt conformance when the topic is something other than name calling.
- diamond559 1y agoIt's bc they are programmed to be agreeable and friendly so that you'll keep using them.
- make3 1y agoisn't that just instruction fine tuning and rlhf inducing style & deference? why is that surprising
- kasey_junk 1y agoWhat’s irritating is that the llms haven’t learned this as bout themselves yet. If you ask an llm to improve its instructions those sort of improvements are what it will suggest. It is the thing I find most irritating about working with llms and agents. They seem forever a generation behind in capabilities that are self referential.
- jstummbillig 1y agoHow are skills different from tools? Looks like another layer of abstraction. What for?
- cynicalsecurity 1y agoSuperpower: AI slop.
- echelon 1y agoI'm sure the horse whip manufacturers had similar things to say about steam powered horses. We just don't think about them much anymore. The whole world is changing around us and nothing is secure. I would not gamble that the market for our engineering careers is safe with so much disruption happening. Tools like Lovable are going to put lots of pressure on technical web designers. Business processes may conform to the new shape and channels for information delivery, causing more consolidation and less duplication. Or perhaps the barrier to entry for new engineers, in a worldwide marketplace, lowers dramatically. We have accessible new tools to teach, new tools to translate, new tools to coordinate... And that's just the bear case where nothing improves from what we have today.
- yoyohello13 1y agoNice try Jensen.
- 4b11b4 1y agoI'm not sure exactly what I just read... Is this just someone who has tingly feelings about Claude reiterating stuff back to them? cuz that's what an LLM does/can do
- apwell23 1y ago[flagged]
- simonw 1y agoHere's a counter-example for you from the another day: https://simonwillison.net/2025/Oct/8/claude-datasette-plugins/ https://simonwillison.net/2025/Oct/8/claude-datasette-plugin... > This isn’t necessarily surprising, but it’s worth noting anyway. Claude Sonnet 4.5 is capable of building a full Datasette plugin now. I do worry a bit about how often I use positive adjectives. If something isn't notable I won't write about it though. In this particle case Jesse's prompting / skills stuff really does deserve the superlatives IMO.
- apwell23 1y agowell explain why OPost is "wild" and what makes you recommend it "strongly" . what have u built with to come to those conclusions ? is this too much to ask.
- simonw 1y agoI recommend it strongly because the "skills" mechanism it describes is a new and very promising technique, and this is the best article I've seen that explains that. It's "wild" because, among many other experiments, Jesse has experimented with giving Claude a "feelings journal" and prompting it using Graphviz DOT diagrams. For my previous writing and work on this you can consult my blog - here's the AI-assisted programming tag: https://simonwillison.net/tags/ai-assisted-programming/ https://simonwillison.net/tags/ai-assisted-programming/
- intended 1y agoThis isnt science, or engineering. This is voodoo. It likely works - but knowing that YAGNI is a thing, means at some level you are invoking a cultural touchstone for a very specific group of humans. Edit - I dug into the superpowers and skills for a bit. Definitely learned from it. There’s stuff that doesn’t make sense to me on a conceptual basis. For example in the skill to preserve productive tensions. There’s a part that goes : > The trade-off is real and won't disappear with clever engineering There’s no dimension for “valid” or prediction for tradeoff. I can guess that if the preceding context already outlines tradeoffs clearly, or somehow encodes that there is no clever solution that threads the needle - then this section can work. Just imagining what dimensions must be encoding some of this suggests that it’s … it won’t work for situations where the example wasn’t already encoded in the training. (Not sure how to phrase it)
- clusterhacks 1y ago> This isnt science, or engineering. > This is voodoo. I was struggling to find the exact reason this type of article bugs me so much, and I think "voodoo" is precisely the correct phrase to sum up my feelings. I don't mean that as a judgement on the utility of LLMs or that reading about what different users have tried out to increase that utility isn't valuable. But if someone asked me how to most effectively get started with coding agents, my instinct is to answer (a) carefully and (b) probably every approach works somewhat.
- theptip 1y ago> some of the ones I've played with come from telling Claude "Here's my copy of programming book. Please read the book and pull out reusable skills that weren't obvious to you before you started reading This is actually a really cool idea. I think a lot of the good scaffolding right now is things like “use TDD” bit if you link citations to the book, then it can perhaps extract more relevant wisdom and context (just like I would by reading the book), weather than using the generic averaged interpretation of TDD derived from the internet. I do like the idea of giving your Claude a reading list and some spare tokens on the weekend where you’re not working, and having it explore new ideas and techniques to bring back to your common CLAUDE.md.
- zahlman 1y ago> It also bakes in the brainstorm -> plan -> implement workflow I've already written about. The biggest change is that you no longer need to run a command or paste in a prompt. If Claude thinks you're trying to start a project or task, it should default into talking through a plan with you before it starts down the path of implementation. ... So, we're refactoring the process of prompting? > As Claude and I build new skills, one of the things I ask it to do is to "test" the skills on a set of subagents to ensure that the skills were comprehensible, complete, and that the subagents would comply with them. (Claude now thinks of this as TDD for skills and uses its RED/GREEN TDD skill as part of the skill creation skill.) > The first time we played this game, Claude told me that the subagents had gotten a perfect score. After a bit of prodding, I discovered that Claude was quizzing the subagents like they were on a gameshow. This was less than useful. I asked to switch to realistic scenarios that put pressure on the agents, to better simulate what they might actually do. ... and debugging it? ... How many other basic techniques of SWEng will be rediscovered for the English programming language?
- hoechst 1y agodocuments like https://github.com/obra/superpowers/blob/main/skills/testing/test-driven-development/SKILL.md https://github.com/obra/superpowers/blob/main/skills/testing... are very confusing to read as a human. "skills" in this project generally don't seem to follow set format and just look like what you would get when prompting an LLM to "write a markdown doc that step by step describes how to do X" (which is what actually happened according to the blog post). idk, but if you already assume that the LLM knows what TDD is (it probably ingested ~100 whole books about it), why are we feeding a short (and imo confusing) version of that back to it before the actual prompt? i feel like a lot of projects like this that are supposed to give LLMs "superpowers" or whatever by prompt engineering are operating on the wrong assumption that LLMs are self-learning and can be made 10x smarter just by adding a bit of magic text that the LLM itself produced before the actual prompt. ofc context matters and if i have a repetitive tasks, i write down my constraints and requirements and paste that in before every prompt that fits this task. but that's just part of the specific context of what i'm trying to do. it's not giving the LLM superpowers, it's just providing context. i've read a few posts like this now, but what i am always missing is actual examples of how it produces objectively better results compared to just prompting without the whole "you have skill X" thing.
- Footprint0521 1y agoI fully agree. I’ve been running codex with GPT Pro (5o-codex-high) for a few weeks now, and it really just boils down to context. I’ve found the most helpful things for me is just voice to Whisper to LLMs, managing token usage effectively and restarting chats when necessary, and giving it quantified ways to check when its work is done (say, AI-Unit-Tests with apis or playwright tests.) Also, every file I own is markdown haha. And obviously having different AI chats for specialized tasks (the way the math works on these models makes this have much better results!) All of this has allowed me to still be in the PM role like he said, but without burning down a needless forest on having it reevaluate things in its training set lol. But why would we go back to vendor lock in with Claude? Not to mention how much more powerful 5o-codex-high is, it’s not even close The good thing about what he said is getting AI to work with AI, I have found this to be incredibly useful in promoting, and segmenting out roles
- 1y ago
- d_sem 1y agoThis article left me wishing it was "How I'm using coding agents to do <x> task better" I've been exploring AI for two years now. It's certainly upgraded itself from the toy classification to a basic utility. However, I increasingly run into its limitations and find reverting to pre-LLM ways of working more robust, faster, and more mentally sustainable. Does someone have concrete examples of integrating LLM in a workflow that pushes state-of-the-art development practices & value creation further?
- jvanderbot 1y agoMy impression is we're still in the tinkering phase. The metrics are coming.
- aydyn 1y agoWhat metrics? We could never objectively measure productivity except in the macro economic sense, so what makes you think we'll be able to now?
- jvanderbot 1y agoThere are best practices of a kind, and well known org structures that work for building software, at least as much as anything can. We'll have some best practice experience with LLM agents soon enough.
- simonw 1y agoMitchell's post from this morning: https://mitchellh.com/writing/non-trivial-vibing https://mitchellh.com/writing/non-trivial-vibing
- 3eb7988a1663 1y agoI am only on the first page and saw this blurb and was immediately annoyed. @/Users/jesse/.claude/plugins/cache/Superpowers/... The XDG spec has been out for decades now. Why are new applications still polluting my HOME? Also seems weird that real data would be put under a cache/ location, but whatever.
- simonw 1y agoIt's in the cache location because it's a copy of a plugin that was installed from a GitHub repository, so that's not the original point of truth for that file.
- wbradley 1y agoI think the point is that ~/.claude should be dispersed among ~/.config/claude, ~/.local/state/claude, etc I agree with this, it’s frustrating that in 2025 apps are still polluting my home dir.
- kibwen 1y agoIt's one thing to wish that apps would put their data anywhere except dumping it in your home dir, but this is exactly why I hate the XDG spec. I want all data for a program--be it the configuration or the cache or the binary itself--to be in a single directory such that 1) "uninstalling" the program, completely and in isolation, is nothing more than just deleting that single directory, and 2) any program not doing arbitrary file I/O can entirely function while having access to only its installation directory, and nothing else on the filesystem.
- 0x6c6f6c 1y agoThis approach couples together everything though, in such a way there's no standard manner of wiping cache but not your app, configuration, etc. XDG may not be perfect but wiping related data for apps following it is straightforward. There are a few directories to delete instead of 1, but still consistently structured at least.
- lcnPylGDnU4H9OF 1y agoThe "How to create skills" link is broken. This is the new location: https://github.com/obra/superpowers/blob/personal-superpowers/skills/meta/writing-skills/SKILL.md https://github.com/obra/superpowers/blob/personal-superpower...
- yoyohello13 1y agoThe post reads like the someone throwing bones and reading their fortune. That part where Claude did its own journaling was so cringe it was hilarious. The tone of the journal entry was exactly like the blog author, which suggests to me Claude is reflecting back what the author wants to hear. I feel like Jesse is consumed in a tornado of llm sycophancy.
- saaaaaam 1y agoClaude has never once said “oh shit” or “holy crap” to me. I must be doing something horribly wrong.
- titanomachy 1y agoYou need to read more books on influencing people. /s
- preommr 1y ago> It made sense to me that the persuasion principles I learned in Robert Cialdini's Influence would work when applied to LLMs. And I was pleased that they did. No, no. Stop. What is this? What're we doing here? This goes past developping with AI into something completely different. Just because AI coding is a radical shift doesn't mean everything has changed. There needs to be some semblance of structure and design. Instead what we're getting is straight up vodoo nonsense.
- imiric 1y ago> Instead what we're getting is straight up vodoo nonsense. It always has been. Starting with the term "AI" itself. Articles like these read the same way to me as any OpenAI announcement from the past 5 years. A bunch of technical mumbo jumbo laced with hyperbole, grand promises of how the technology is changing the world, and similar platitudes. I've learned to filter most of it out. Occasionally I'll stumble upon an actually useful and practical tidbit of information which I can apply in my own workflow, which does involve LLMs, but most of the time it's just noise.
- w10-1 1y ago> what we're getting is straight up voodoo nonsense Maybe not in this case. For the AI to create a solution, it has to come up with a vector for your intention and goals. It makes some sense for an AI trained on human persuasion materials (basically, everything has a rhetorical aspect) to also track human persuasion features for intentions. However, results will vary. Just as people trying to deploy rhetorical techniques (and ridiculous power stances) often come off as foolish, I believe trying to hack your intention vector with all-caps and super-superlatives won't always work as intended (pun intended). Still, if you find yourself not getting what you want, and you check your prompt and find some persuasion feature missing (e.g., authority), I think it's worth trying to add something on point.
- Fargren 1y ago> It makes some sense for an AI trained on human persuasion Why? > However, results will vary. Like in voodoo? I'm sorry to be dismissive, but your comment is entirely dismissing the point it's replying to, without any explanation as to why it's wrong. "You are holding it wrong" is not a cogent (or respectful) response to "we need to understand how our tools work to do engineering".
- imiric 1y agoAnd here I am in October 2025 still using "AI" tools via a chat UI in Emacs, like a caveman. I've written some code to help me with managing context and such, but the tools are there when I need them, and otherwise stay out of my way. I have no interest in trying to understand the thought process of people who write and work like this. They're more interested in chasing the latest overhyped trends produced by tech companies and influencers, than actually producing quality software that solves real-world problems. It's some weird product of the tech and social media echo chambers they perpetually live in, which I find difficult to describe. But apparently I have to learn about "skills" and "superpowers" now... Give me a break.
- JaggerFoo 1y agoI don't see any code. Where are the examples of use on real code?
- orangebread 1y ago[dead]
- meander_water 1y agoThe problem with stuff like this is that it's hard to evaluate. You don't even know when the agent is using a skill, or if the skill even made a difference. Using tools lets you at least instrument tool calls, and control what gets executed.
- redhale 1y agoI agree, I think traceability will be extremely important in evolving and improving a system like this. Since scripting is involved in searching for and managing skills, I feel like there is probably a way to achieve some kind of use tracing, but I'm not quite sure. Seems like this, if implemented, could also be fed back into the system for self improvement.
- deleted 1y ago[deleted]
- herval 1y agoFascinating write-up. I loved this bit of debugging: > The first time we played this game, Claude told me that the subagents had gotten a perfect score. After a bit of prodding, I discovered that Claude was quizzing the subagents like they were on a gameshow. This was less than useful. I asked to switch to realistic scenarios that put pressure on the agents, to better simulate what they might actually do. Also his Claude says shit a lot
- Aloisius 1y agoWhat's up with people (or I suppose AI) including copyright licenses in AI generated code? At least it's an MIT license, but since AI output isn't copyrightable, I'm unsure what the point is since people can legally ignore the license.
- hugh-avherald 1y ago^ (not legal advice -- far from it)
- Aloisius 1y agoIf there some reason why one wouldn't be able to ignore the copyright license of something not protected by copyright, I'd love to hear it. The copyright office has been quite clear (rightly so imo) that AI output is not protected by copyright without substantial human creative expression in the final product and purely prompt-created works simply don't qualify. Indeed, I expect people muddling their codebases with AI output are going to find themselves in an interesting position of having to prove how much code humans actually wrote to enforce copyright claims if their code ever gets leaked.
- kragen 1y agoThat's just the copyright office of one country out of a couple hundred, the courts can overrule them, and legislation can change. However, I agree that currently in the US (or on code written in the US) copyright probably doesn't inhere in AI-written code.
- Aloisius 1y agoThe US constitution limits copyright to protection for authors and inventors. I'm skeptical that a simple law could extend protection to machine generated works without being ruled unconstitutional nor does there appear to be any significant government or public support for such a thing. And while yes, the US is just one country, but it does have a bit of an outsized software development industry. I also haven't hear of any other countries lining up to give machine-generated works copyright protection.
- benrutter 1y agoI'm so curious around what people's median experience is of AI coding tools. I've tried agents every now and then, recently for something very simple- add an option to request csb format in a data api. The results were, well, not good. . . I ended up undoing literally all changes because writing from scratch was a lot easier than trying to refactor the total mess it has made from what I'd have thought was a trivial feature. I haven't done loads of prompt engineering etc, in all honesty it seems a lot of work when I haven't seen promise yet in the tool. I see articles like this, and I always wonder, am I the outlier or is the writer? My experience of agentic AI is so hugely different to what some people are finding.
- x0x0 1y agoI'm the same, with the same question if it's me. I've had success with eg spitting out templated html; sometimes with css; sometimes with writing tests where I'm very specific about what I want (set up these structures, test this condition), etc. It's mediocre (good start, very far from production) with writing screens in react native. It does slightly better on rails, but far from production ready. After that, it kinda works, but my effort level to turn the output into working code is higher than just writing it myself.
- aydyn 1y agoThink of this: whats the likelihood that what you are asking for would be found in some public github repo? If its high then you are good to go.
- a123b456c 1y agoI think you're pointing in the right direction, but I would rephrase as, what's the likelihood that the solution exists in the github repo in a way that the machine can recognize as relevant to your prompt? If many versions of the solution exist, due to the problem's common occurrence, and if you can evaluate the LLM's output, then you're good to go.
- cosmodust 1y agoIt's very use case specific, I find them really good in simple repetitive tasks as long as you guide them at low level. Although you do need to keep a close eye as they easily spoil your existing work.
- kreyenborgi 1y agoAnyone else get the feeling like CLAUDE.md fiddling is the new dotemacs fiddling?
- zkmon 1y ago<Homer Simpson mode>Oh yeah? If prompting is such damn cool hard thing, why can't I ask my AI slave to do all this prompting mumbo jumbo for me?</Homer Simpson mode>
- tobbe2064 1y agoIs it possible to set up this kind of workflow with the plug in that comes bundled with vs code, given that you have an enterprise github copilot account that includes Claude?
- redhale 1y agoSubagents are a critical feature that GH Copilot still lacks. They allow your main agent to use another agent as a tool, meaning the main agent's context doesn't get nearly as polluted. Good read on the benefits of this pattern: https://jxnl.co/writing/2025/08/29/context-engineering-slash-commands-subagents/ https://jxnl.co/writing/2025/08/29/context-engineering-slash...
- d4rkp4ttern 1y agoA big issue working with code agents is what I call context-recall: restoring context when working on a new feature or fix, that builds on recent work. Meaning, the previous work may have involved multiple CLI sessions, summaries dumped to various markdown files like documentation files, plan files, issue files, PR-descriptions etc. Then when starting new work with a code agent you have to hunt down all of this scattered context from various md files and session logs to fill in background for the code-agent about what was recently done. I see many workflows that help with working on a fresh feature or fix, but nothing that addresses context-recall. But maybe the OP workflow or others do that, I haven’t dug too deep into them.
- d4rkp4ttern 1y ago(Just released the OP blog actually does address exactly this)
- dwb 1y agoHonestly, if the LLM/agent can't do what I want with a simple, shortish prompt that I understand, augmented by some well-chosen tool calls, I'm not interested. These incantations may or may not work, but I just don't want them. Reams of vague twiddling of an unknowable black box. I want the amount of mystery kept at an absolute minimum when I'm programming.
- iamjfu 1y agoI am interested by this link: https://blog.fsck.com/blog/2025/superpowers/superpowers-demo.txt https://blog.fsck.com/blog/2025/superpowers/superpowers-demo... ``` Claude Code v2.0.13 Sonnet 4.5 (with 1M token context) Claude Max /Users/jesse/tmp/new-tool/.worktrees/todo-cli ``` How does this person have access to Sonnet 4.5 with 1m token context? I don't see this referenced anywhere when I search or when I ask Claude about it.
- d4rkp4ttern 1y agoIt’s a limited release beta feature not available to all. You can try to activate it by doing: /model sonnet[1m] And it accepts it but the at the next API call it may fail and say “this beta model is not available with your subscription”. I haven’t gotten access yet. One of the nice things about Codex (GPT-5) is the supposed 400k token context (although performance starts to deteriorate when you get to 80% context usage).
- piperswe 1y agoOpenRouter shows Sonnet 4.5 as having a 1M context limit: https://openrouter.ai/anthropic/claude-sonnet-4.5 https://openrouter.ai/anthropic/claude-sonnet-4.5
- AlexCoventry 1y agoI think this is cool, but some performance benchmarks would really help to sell it.
- throw-10-13 1y ago“Here is a collection of arcane incantations and humiliating prostrations I use to get my AI homunculus to serve me.” Having to beg and emotionally manipulate an agent into doing what you want goes so far beyond black-box that I find it difficult to believe these people actually get useful work done using these tools. I generally consider myself pro-ai in the workplace, but this nonsense is starting to change my mind.
- StapleHorse 1y agoA little bit off topic. I love how AI is advancing so fast that the usual title: "How i'm using XX in 20NN" is not specific enough, now we need the month.
- novoreorx 1y agoTo me, this kind of stuff is like bloated boilerplates such as "full-stack e-commerce SaaS NextJS boilerplate." I never use them because I want more control and fewer unpredictabilities. They seem to save you some time, but you will pay a lot more for it later when you encounter deep bugs or need to refactor. For this reason, I won't use prompt templates for agentic coding tools either. There have been enough suggestions to write your own AGENTS.md and not overcomplicate the prompts.