4 ms·
We really need some consolidation around commands, skills, subagents, and plugins. For example, if you want to, say, review code, you have five options now: -
by mil22 4mo ago
We really need some consolidation around commands, skills, subagents, and plugins. For example, if you want to, say, review code, you have five options now:
- Write a .claude/commands/review.md. Simple but deprecated.
- Use a /code-review skill, either one you install or one you just write yourself (it's just Markdown, after all).
- Use the /pr-review subagent. Also just Markdown, but it runs "in the background" and "in parallel", so it must be better, I guess.
- Install the /code-review plugin. This just installs the skills and subagents above.
- Simply ask Claude to review the code. Probably works almost as well as the above in most situations.
They are all just variations of "insert a canned prompt", varying only along the dimensions of (a) how and where the prompt is installed and from where it is sourced, and (b) which context or contexts the prompt runs in. There's not much advice here about which option is best, and no clear best practices seem to have emerged yet either. Personally, I find just asking Claude to review the code works well enough.
Some of the advice here is also off. For example:
"Install a language server plugin. Type errors and unused imports caught after every edit. Highest-impact plugin you can install."
I work mostly with Rust, Python, and Dart, and followed similar advice, installing LSPs for all three in both Claude Code and Codex. Two months later, after heavy development in all three languages and hundreds of sessions - and frequently running out of RAM due to all the Rust analyzer, Dart analysis server, and Ty LSP servers the harnesses were spinning up - I checked the session logs to see how often the agents were actually invoking the LSP tools. The answer was they had invoked them literally once the entire time. I uninstalled all my LSPs and haven't looked back. The agents do just fine using ripgrep and calling cargo clippy, dart analyze, ty check, etc. themselves.
- para_parolu 4mo agoI just consider this temp phase because models are dumb and harnesses are not yet there. When I need code review I should just say “review it”. Model should figure out what plugins, skills, etc. to use.
- sheept 4mo agoWhy does it need plugins/skills for a code review? Claude will just "review it" if you ask it to, and if you have particular preferences, they can go in CLAUDE.md
- unshavedyak 4mo agoSkills are effectively the same thing as asking it, just with more depth. So the skill is just a framework for a very precisely asked question. It often includes how you want Claude to respond, etc. I’m not aware of anything fundamentally unique about skills or commands, they’re just more tokens to shape the llm
- bcherny 4mo agoTotally. You can do that now, and Claude will know to use /code-review.
- Izmaki 4mo agoI imagine that the companies that earn money from input and output tokens really, really like excessive skills because of the sheer amount of potentially pointless constraints and instructions being sent back and forth ("don't store passwords as plaintext", "always check for syntax errors" and other obvious guidelines).
- cheema33 4mo agoMy personal experience is the opposite. Lack of skills uses more tokens.
- nlawalker 4mo ago> They are all just variations of "insert a canned prompt", varying only along the dimensions of (a) how and where the prompt is installed and from where it is sourced, and (b) which context or contexts the prompt runs in. Yes, yes, thank you, sometimes I feel like I'm taking crazy pills. The industry and overall developer ecosystem has become absolutely mesmerized by the act of creating and popularizing little bits of protocol and machinery to dress up the act of inserting text into the machine. Yes, they're useful and provide some consistency, but I'm convinced that the main reason people like them so much is because they put a thin "I'm still a programmer wielding complicated tools that laypeople don't understand" coating over the fact that we're all just asking the AI nicely to do a thing.
- bcherny 4mo agoHey, Boris from the CC team here. I agree, we're working on consolidating these. Going forward it will just be the built-in /code-review skill. Here's how to use the skill on the latest version: /code-review # do a balanced code review. checks for bugs and inconsistencies, poor code quality, duplication, band aids, etc. /code-review --fix # same as above, but also fix the issues # choose an explicit effort level (defaults to your current effort level). all of these also accept --fix: /code-review low /code-review medium /code-review high /code-review xhigh /code-review max # do an expensive and extremely thorough review (reliably catches >99% of bugs, costs $3-20 per review depending on complexity): /code-review ultra Open to feedback if anyone has feedback or ideas for how to make these even nicer to use.
- extr 4mo agoHey Boris, some feedback. I like the new /code-review skill but was disappointed you guys removed /simplify because I quite liked the focus on finding code reuse/efficiency opportunities. I see now in 2.1.152 you added those focus areas back to /code-review, but still bundled with the correctness finding. It would be great to have more fine grained control over the /code-review angles beyond just effort level. Or maybe you would recommend that I just specify that as freeform input after effort level?
- bix6 4mo agoHi Boris, what is the advantage of using /code-review vs just asking Opus to “code review”? As a casual user working on hobby projects, I struggle to keep up with the pace of changes and knowing what to use when. My default now is to use Opus for all coding (sonnet is fine but seems dumber) and to prompt it for everything I need. I’ve had great success with this but clearly I’m missing power user functions with the slash commands and such.
- Majromax 4mo ago> They are all just variations of "insert a canned prompt", varying only along the dimensions of (a) how and where the prompt is installed and from where it is sourced, and (b) which context or contexts the prompt runs in. There's not much advice here about which option is best, and no clear best practices seem to have emerged yet either. Personally, I find just asking Claude to review the code works well enough. The subagent approach is structurally different from the others because it runs with clean context. That has three major effects: 1. All other things being equal, it will result in a lower cost-to-solution because of the quadratic cost scaling of an LLM session (input token or cached-input cost being paid with each new round). 2. The review model will not be able to 'cheat' by retaining assumptions from the main session, such as "x must be done like y." For people, this is why having a separate person perform code review (or, if not possible, reviewing code after a mind-clearing break) is handy; the applicability of this analogy to LLMs is vague but reasonable. 3. The main model will only see the results of the review, not the detailed reasoning that leads up to it. On one hand this avoids more context pollution, but on the other hand it might lead to duplicative logic to re-discover the mechanics behind bugs found. > I checked the session logs to see how often the agents were actually invoking the LSP tools. The answer was they had invoked them literally once the entire time. I think the intent behind 'install a language server plugin' is that these tools should lint automatically after every edit, without waiting for an explicit call from the LLM.
- mil22 4mo ago> The subagent approach is structurally different from the others because it runs with clean context. Yes, and this is what I mean by "which context the prompt runs in". The subagent approach is different and has pros and cons, and it may in some situations be better (but perhaps not in others). On the other hand, I can also just create a new conversation and paste my own review prompt into it; then take the last turn's summary output and feed it back into my main conversation thread in the unusual event I would need to do so. Spawning a subagent is a convenient shortcut for this, but ultimately, it's the same thing. > I think the intent behind 'install a language server plugin' is that these tools should lint automatically after every edit, without waiting for an explicit call from the LLM. This is a great point and I had only checked my session logs for explicit tool calls. I went back and looked for diagnostics injected automatically by the harness after every edit, and whether the agent made use of them. Claude: neither the Rust or Dart LSPs ever inserted any diagnostic events, but Ty did. Across 627 sessions, ty-lsp injected diagnostics blocks in 186 sessions, with a total of 33 findings. Out of those 33, 32 were dismissed as unrelated (13) or pre-existing (19). Only 1 finding was acted upon. The model is in the habit of running the batch analysis tools (ruff, ty, cargo clippy etc.) and prek anyway, so it would have caught that diagnostic regardless. Codex: no diagnostic events were inserted by any of the LSPs. So I won't be reinstalling those LSPs.
- mdav75 4mo ago[flagged]
- superfrank 4mo agoI have thought for a while now that skills were a bad abstraction. There's a lack of definition around what to use them for that I think contributed to why they rose to the top, but that's also why I think they aren't a good long term option. The fact that I can have a skill that is just general guidance on front end design best practices that an agent can call upon whenever they feel, and another that is essentially a run book of steps that need to be followed exactly only when explicitly triggered, and a third that is basically just instructions on how to use a specific tool and all of those are acceptable just feels wrong to me. I get why it caught on and why the flexibility is attractive when the entire world is collectively learning a new tool, but skills have come to feel like the junk drawer in the kitchen where you just throw random shit when you don't want to think about a better place to put it. I would love to see the world standardize on something like: - Agents: Essentially personalities for a model to take on. This becomes the new place for skills like "front end expert" where you're not telling an agent to do a specific thing, just to think in a certain way about a task. - Prompts: Repeatable instructions for specific tasks that an agent should follow when prompted. This could be something like a checklist style run book on how to resolve a certain error that an agent needs to follow exactly or it could be something like here's an idea I have for a new feature please poke holes in it. - Tools: Tools (like CLIs, MCPs, or scripts) and instructions on how and when to use them. I'm purposefully not calling this skills because I think the term is overloaded, but that's kind of what this is.
- binarymax 4mo agoI mostly agree here. I just treat skills as “prompts”, and I scope them to domain specific tasks. I’m surprised when I see skill files that are really short. Most of mine are pages long.
- dyauspitr 4mo agoHonestly I don’t like that we’re coming back around to command line terminology you have to know and remember on a natural language intelligence. Codex doesn’t do this crap yet right?
- sfrangulov 4mo ago[flagged]