6 ms·
It seems the benchmarks here are heavily biased towards single-shot explanatory tasks, not agentic loops where code is generated: https://github.com/drona23/cla
by btown 6mo ago
It seems the benchmarks here are heavily biased towards single-shot explanatory tasks, not agentic loops where code is generated: https://github.com/drona23/claude-token-efficient/blob/main/BENCHMARK.md https://github.com/drona23/claude-token-efficient/blob/main/...
And I think this raises a really important question. When you're deep into a project that's iterating on a live codebase, does Claude's default verbosity, where it's allowed to expound on why it's doing what it's doing when it's writing massive files, allow the session to remain more coherent and focused as context size grows? And in doing so, does it save overall tokens by making better, more grounded decisions?
The original link here has one rule that says: "No redundant context. Do not repeat information already established in the session." To me, I want more of that. That's goal-oriented quasi-reasoning tokens that I do want it to emit, visualize, and use, that very possibly keep it from getting "lost in the sauce."
By all means, use this in environments where output tokens are expensive, and you're processing lots of data in parallel. But I'm not sure there's good data on this approach being effective for agentic coding.
- sillysaurusx 6mo agoI wrote a skill called /handoff. Whenever a session is nearing a compaction limit or has served its usefulness, it generates and commits a markdown file explaining everything it did or talked about. It’s called /handoff because you do it before a compaction. (“Isn’t that what compaction is for?” Yes, but those go away. This is like a permanent record of compacted sessions.) I don’t know if it helps maintain long term coherency, but my sessions do occasionally reference those docs. More than that, it’s an excellent “daily report” type system where you can give visibility to your manager (and your future self) on what you did and why. Point being, it might be better to distill that long term cohesion into a verbose markdown file, so that you and your future sessions can read it as needed. A lot of the context is trying stuff and figuring out the problem to solve, which can be documented much more concisely than wanting it to fill up your context window. EDIT: Someone asked for installation steps, so I posted it here: https://news.ycombinator.com/item?id=47581936 https://news.ycombinator.com/item?id=47581936
- david_allison 6mo agoIs this available online? I'd love documentation of my prompts.
- sillysaurusx 6mo agoI’ll post it here, one minute. Ok, here you go: https://gist.github.com/shawwn/56d9f2e3f8f662825c977e6e5d0bfc08 https://gist.github.com/shawwn/56d9f2e3f8f662825c977e6e5d0bf... Installation steps: - In your project, download https://gist.github.com/shawwn/56d9f2e3f8f662825c977e6e5d0bfc08 https://gist.github.com/shawwn/56d9f2e3f8f662825c977e6e5d0bf... into .claude/commands/handoff.md - In your project's CLAUDE.md file, put "Read `docs/agents/handoff/*.md` for context." Usage: - Whenever you've finished a feature, done a coherent "thing", or otherwise want to document all the stuff that's in your current session, type /handoff. It'll generate a file named e.g. docs/agents/handoff/2026-03-30-001-whatever-you-did.md. It'll ask you if you like the name, and you can say "yes" or "yes, and make sure you go into detail about X" or whatever else you want the handoff to specifically include info about. - Optionally, type "/rename 2026-03-23-001-whatever-you-did" into claude, followed by "/exit" and then "claude" to re-open a fresh session. (You can resume the previous session with "claude 2026-03-23-001-whatever-you-did". On the other hand, I've never actually needed to resume a previous session, so you could just ignore this step entirely; just /exit then type claude.) Here's an example so you can see why I like the system. I was working on a little blockchain visualizer. At the end of the session I typed /handoff, and this was the result: - docs/agents/handoff/2026-03-24-001-brownie-viz-graph-interactivity.md: https://gist.github.com/shawwn/29ed856d020a0131830aec6b3bc29dc9 https://gist.github.com/shawwn/29ed856d020a0131830aec6b3bc29... The filename convention stuff was just personal preference. You can tell it to store the docs however you want to. I just like date-prefixed names because it gives a nice history of what I've done. https://github.com/user-attachments/assets/5a79b929-49ee-4614-8f42-e764793ace5e https://github.com/user-attachments/assets/5a79b929-49ee-461... Try to do a /handoff before your conversation gets compacted, not after. The whole point is to be a permanent record of key decisions from your session. Claude's compaction theoretically preserves all of these details, so /handoff will still work after a compaction, but it might not be as detailed as it otherwise would have been.
- hatmanstack 6mo agoSeems crazy to me people aren't already including rules to prevent useless language in their system/project lvl CLAUDE.md. As far as redundancy...it's quite useful according to recent research. Pulled from Gemini 3.1 "two main paradigms: generating redundant reasoning paths (self-consistency) and aggregating outputs from redundant models (ensembling)." Both have fresh papers written about their benefits.
- whattheheckheck 6mo agoNo such thing as junk DNA kinda applies here
- wongarsu 6mo agoThere was also that one paper that had very noticeable benchmark improvements in non-thinking models by just writing the prompt twice. The same paper remarked how thinking models often repeat the relevant parts of the prompt, achieving the same effect. Claude is already pretty light on flourishes in its answers, at least compared to most other SotA models. And for everything else it's not at all obvious to me which parts are useless. And benchmarking it is hard (as evidenced by this thread). I'd rather spend my time on something else
- scosman 6mo agoalso: inference time scaling. Generating more tokens when getting to an answer helps produce better answers. Not all extra tokens help, but optimizing for minimal length when the model was RL'd on task performance seems detrimental.
- joquarky 6mo agoI liked playing with the completion models (davinci 2/3). It was a challenge to arrange a scenario for it to complete in a way that gave me the information I wanted. That was how I realized why the chat interfaces like to start with all that seemingly unnecessary/redundant text. It basically seeds a document/dialogue for it to complete, so if you make it start out terse, then it will be less likely to get the right nuance for the rest of the inference.
- alsetmusic 6mo ago> No explaining what you are about to do. Just do it. Came here for the same reason. I can't calculate how many times this exact section of Claude output let me know that it was doing the wrong thing so I could abort and refine my prompt.
- deleted 6mo ago[deleted]
- deleted 6mo ago[deleted]
- heyethan 6mo ago[flagged]
- dataviz1000 6mo agoI made a test [0] which runs several different configurations against coding tasks from easy to hard. There is a test which it has to pass. Because of temperature, the number of tokens per one shot vary widely with all the different configurations include this one. However, across 30 tests, this does perform worse. [0] https://github.com/adam-s/testing-claude-agent https://github.com/adam-s/testing-claude-agent
- btown 6mo agoThis is an amazing analysis! Thank you for running this :)
- baq 6mo agoif the model gets dumber as its context window is filled, any way of compressing the context in a lossless fashion should give a multiplicative gain in the 50% METR horizon on your tasks as you'll simply get more done before the collapse. (at least in the spherical cow^Wtask model, anyway.)
- matchagaucho 6mo agoSome redundancy also helps to keep a running todo list on the context tip, in the event of compacting or truncation. Distilled mini/nano models need regular reminders about their objectives. As documented by Manus https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus https://manus.im/blog/Context-Engineering-for-AI-Agents-Less...
- 0xbadcafebee 6mo agoThere's an ancient paper that shows repetition improves non-reasoning weights: https://arxiv.org/html/2512.14982v1 https://arxiv.org/html/2512.14982v1
- hrmtst93837 6mo ago[flagged]
- sossov 6mo ago[dead]