6 ms·
The Token Compression Illusion: Why I'm Skeptical of RTK
- breadislove 4mo agoslop complaining about other slop
- lackoftactics 4mo agothank you, author here. I will stay civil here and focus on the rtk as that was the goal of article. So do you think rtk cli is ai slop? I had some suspicions looking at their repo and number of issues and their style. The prettier issue with running successfully while binary wasn't even installed was quite entertaining
- grey-area 4mo agoDid you use an LLM for the blog post? it reads like it in places.
- lackoftactics 4mo agoI have raycast shortcut for fix grammar, it might done more damage than adding a, an, the or changing tenses.
- gowld 4mo agoA content-free 2nd "paragraph" like this turned me off immediately. > But in the current dev tools gold rush, if something sounds too good to be true, it almost always is. The people who are interested in RTK and in criticism of RTK aren't interested in pablum like this.
- lackoftactics 4mo agook, this one is all mine. So that's even more hurtful as this is 100% me
- danr4 4mo agothis is aboslutely entirely written by AI
- lackoftactics 4mo agoAs an author of the text, I can say you are „absolutely” not correct. I might be already spending too much time with llms and they start to shape my texts, so I am not proud of that either. But thanks for bringing very valuable insight to otherwise interesting discussion.
- grey-area 4mo agoI hope you do take this to heart and stop using LLMs in any form for writing, people are not telling you this just to annoy you but because you owe your readers more than that. LLMs generate superficially plausible text, not good writing, use your own voice always.
- citizenpaul 4mo ago> they start to shape my texts I'm a bit curious about this? What draws you to speaking like an LLM rather than your own voice? I've personally never felt compelled to be more llm like in my writing.
- SubiculumCode 4mo agoI feel like what is needed is not compression, but aggressive context management with subagents.
- lackoftactics 4mo agoI am the author the text. What do you mean by aggresive context management with subagents? Would you add a lopp that would trim the context? Both of those tasks seem even more difficult
- skinfaxi 4mo agoI believe they mean aggressive delegation to minimize context bloat in the coordinating agent.
- lackoftactics 4mo agothat would make more sense, trimming context with subagents sounds like an overkill
- svachalek 4mo agoThis is a really useful technique in my experience. The harnesses are starting to do it more on their own but if you encourage the use of more subagents, I find it's typically nothing but win.
- SubiculumCode 4mo agoFirst, I only say this because of what I learned as a phD inhuman memory, not as someone who authors agentic workflows or does AI. How human cognition tends to work by simultaneously utilizing and combining/separating multiple frequency scales of information. A simple way of thinking about is this: We tend to encode and retrieve both the gist of what is happening, and the verbatim details of what happened. The gist can be thought of as low frequency information, almost like bullet points, that contain the big overview goal, keypoints). The verbatim traces, are the high resolution memory that contains all the details. The gist helps encoding and recall by providing encoding and retrieval context cues. There are also levels in between those two, but I was keeping it simple. During human development, verbatim memory capacity increases first, but then hits a wall/plateau. Further performance increases begin to depend on the ability to utilize and gain from gist-like representations that can guide encoding and retrieval of verbatim details within contexts. You don't need to keep everything in the context window. My untested, perhaps naive hypothesis is that what is needed is that sub-agents dealing with verbatim tasks (actually writing code), their context window should be managed by an agent above that is tuned to information at a lower frequency, and it by another above it on even lower frequency information. Lowest frequency information context windows feel up slowly. High-frequency information fills up fast. Use the low frequency information to retrieve the needed high frequency information.
- iam-TJ 4mo agoAm I the only one that thought RTK was Real-Time Kinematics used for precision with satellite navigation?
- dayjaby 4mo agoNo. I clicked here for the same reason.
- lackoftactics 4mo agoI might have picked better title, but they are literally called rtk https://github.com/rtk-ai/rtk https://github.com/rtk-ai/rtk and it stands for Rust Token Killer
- arcanemachiner 4mo agoI've been trying out RTK and it seems kinda alright. I doubt it's saving much, but the quality of the work feels similar. But if it's making a dent in token usage (which I have not personally measured), then that's great. I had to add some system prompt instructions to Pi to help it work (GPT 5.5 initially got confused when `git status` looked different than expected). The Claude Code extension appears to do a proper job of informing the agent about the unexpected shape of the output without any extra work on my part.
- lackoftactics 4mo agoso how do you justify it's usage if it's not saving much and the work feels similiar. They have 664 issues open and some of them are quite funny, the tools are called and return success even though they aren't even installed. My take is that handling so many versions and so many different tools shouldn't be the work of any single repo. The responsibility should be either on coding agent to compress or best case scenario people who are responsible for cli tool
- arcanemachiner 4mo agoI'm not justifying its usage, and I don't have to. I've been trying it out for a couple days and it seems kinda OK or whatever. If that upsets you, then that's your problem. I might dump it later on if it doesn't provide much if a benefit. I typically try out new things, then cull whatever doesn't work. This tool seems pretty neutral for now, at least.
- lackoftactics 4mo agono, it doesn't upset me. I am open for discussion, there might be things I miss and don't understand. I am just trying to get why it's been pushed so hard lately and if the benefits are really there. Sorry, if I sounded upset to you, but I am trying to be really civil and just genereally curious
- arcanemachiner 4mo ago
- old_sysadmin 4mo agoI feel like the state of the art is baked into the compaction logic, and I've had a lot of problems with compaction (absent other prompting) losing key bits of state. https://github.com/toon-format/toon https://github.com/toon-format/toon is another interesting one, and I feel like it takes on a much more achievable goal - reduce whitespace and verbosity of JSON, not overall context compression.
- arcanemachiner 4mo agoPersonally, I find compaction to be unreliable, which forces me to rely heavily on session-specific planning documents and inter-agent handoff messages.
- KingMob 4mo agoUnfortunately, I saw some research paper suggesting that TOON doesn't help overall. It's shorter than JSON, but novel/different enough that agents wasted more tokens thinking about it, interpreting it, and making extra calls.
- compuficial 4mo ago> 1. Gamified Savings vs. Your Actual API Bill Tool use output represents a large amount of my output. I'll take 3.7M tokens saved on 3.9M tokens of input. Tokens saved are tokens saved. > 3. Where Are the Accuracy Benchmarks? As a user of RTK, it would be nice to see accuracy benchmarks. However, I've seen no evidence of the model missing anything critical as a result of the compression. As part of their design philosophy they are very strict about preserving correctness to the point that if a filter fails they fall back to raw output. For my most frequently used commands I've inspected the source, was happy with what I saw, they've earned my trust thus far. > The day git, cargo, npm, or grep updates its terminal formatting by a few spaces or changes an error layout, RTK's regex and parsing filters will break. And returning to the silent failure trap, it won't throw an explicit error; it will fail quietly, feeding corrupted or partial text to your agent. Again, any filter that fails simply falls back to the raw output. One of their core pillars is avoiding this exact scenario you described. RTK should never feed corrupted or partial text to an agent. Your concerns are fair but I'd like to see your criticism backed up with evidence. Have you used RTK? Have you found evidence that they are failing to preserve correctness?
- lackoftactics 4mo agoI was looking through the issues as investigation. Some issues that caught my attention are looking quite bad https://github.com/rtk-ai/rtk/issues/2494 https://github.com/rtk-ai/rtk/issues/2494 https://github.com/rtk-ai/rtk/issues/2462 https://github.com/rtk-ai/rtk/issues/2462 https://github.com/rtk-ai/rtk/issues/2395 https://github.com/rtk-ai/rtk/issues/2395
- compuficial 4mo agoFwiw, I just ran the steps to reproduce and got `Error: prettier produced no output` on rtk (0.42.2). Not saying this isn't valid for the users environment but I could not reproduce on linux.
- lackoftactics 4mo agoappreciate the engineering effort and skin in the game. I might try on macos today as the author of issue.
- cityofdelusion 4mo agoI am glad articles like this are finally starting to get some momentum around what I call the LLM magic box industry. From caveman mode to RTK to semantic search and everything in between. Developers have become magicians that cast spells instead of engineers. It sucks at work especially with everyone so sure that their magic spell is the one for ultimate token savings. My criteria are: if it’s not in a harness it’s probably not that good (the best ideas float up to Codex/Claude imo) and any GitHub advertising some percent of token savings is not to be trusted. It’s hard to avoid the snake oil and I hope people start thinking critically on this stuff.
- blubber 4mo agoThere is a conflict of interest, though.
- baq 4mo agoOnly in inference, but if you consider that they’re reinvesting inference performance in training I think the conflict argument is overblown.
- arcanemachiner 4mo agoThe idea itself is sound: If you can reduce the signal-to-noise ratio in the context window, then that's a good thing. Whether or not RTK actually does this has not been established. I would be glad to see some proper benchmarks done on the actual difference this tool makes (not some meaningless "up to 90%" type of language).
- lackoftactics 4mo agoI was wondering if that impacts the accuracy, obviously the rtk output wasn't in the training dataset, but maybe it doesn't matter at the end
- celrod 4mo agoI found this, which has some: https://arxiv.org/pdf/2605.28876 https://arxiv.org/pdf/2605.28876 TLDR: RTK does not look good according to the author's benchmark.
- Catloafdev 4mo agoI don't agree with the conclusion at all. I can see the value of RTK - whether it is buggy or vibe coded is kind of secondary. That basically comes down to how severe and often the bugs are. There's no gamification of savings here. Tool output can be meaty. Is the author skeptical of the concept, or the implementation? Because only one of those is worth critiquing.
- lackoftactics 4mo agoHey, author here, I am skeptical of implementation starting from Rust Token Killer and looking to monetize on Rust love by other developers. Concept is fine to me and I believe we should optimize, but a repo that will handle all tools sounds like Sisyphus rolling a rock up the hill.
- tlarkworthy 4mo agoI tried it and it does not compress messages which was 90% of my context, so it only compresses a small part of my token usage. If you read it carefully you will realize that is exactly stated. If you look at /context you will probably see that tool calls are not where you are spending token on, so a proxy that compresses tool calls will not make much impact, whilst still being true that it compresses tool calls by 8x. Its just not that important for long coding sessions for me. "native/built-in Read or cat tools, the data is not intercepted by RTK's shell hook"
- lackoftactics 4mo agoAuthor of the text here. I will be honest with why I wrote it, the rtk ai looks very odd to me as software engineer, the number of stars, no mention of accuracy and how management is pushing that stuff to optimize costs. Now people are wrapping every possible command in rtk and trying to handle all major possible commands and decide which output you should get.
- ianwalter 4mo agoWhy didn’t you offer any real world usage numbers to illustrate your point? I found this unhelpful.
- lloyd-christmas 4mo agoI read another post oddly similar earlier today that has more explicit data on that authors codebase: https://codepointer.substack.com/p/cutting-llm-token-costs-with-rtk https://codepointer.substack.com/p/cutting-llm-token-costs-w... TLDR; ~3-4% savings to actual API costs with rtk, caveman, and headroom combined, but nothing tangible on if those cost reductions came at a cost of quality. By their calculations, rtk saved them $4.96 on a $926 bill.
- bcollins34 4mo ago^recommend reading this one
- lackoftactics 4mo agoThat’s the fair point. The rtk promotional posts point to 60-90% tokens savings and there is no mention how they perform accuracy wise. The commenter below did great job pointing to resource showing caveman, rtk saving just couple bucks on $926 bill. Thanks, Llyoyd Christmas for linking to useful substack
- fumeux_fume 4mo agohttps://en.wikipedia.org/wiki/Brandolini%27s_law https://en.wikipedia.org/wiki/Brandolini%27s_law
- blubber 4mo ago"Where Are the Accuracy Benchmarks?" I wish the author would have provided one.
- graphememes 4mo agoI don't disagree with the article, but I also don't disagree with RTK. The output of these commands is not optimized for agents (or humans) for that matter.
- trjordan 4mo agoThe core of the problem is that there are a million tools that make AI better, and no ways to measure whether AI is working better. Big companies with popular products have it. They do something between normal product analytics and chatbot evals to figure out if users are being successful in their sessions. That's the job. But any given dev, with between 3 and 50 sessions a day? Like, I have no idea what makes the LLM better. It's all vibes. My company has a whole stack here. Preferred harnesses, preferred models, skills, the shape of our code, everything. There's gotta be a way to measure whether this setup is working for us, at 1 / 1-million-th the scale of a Claude Code.
- jahala 4mo agoThere is an answer- these tools should benchmark by cost per correct answer - not just tokens saved.
- lackoftactics 4mo agoAnd the effort to produce valid benchmarks is tremendous. You are probably right and that’s very annoying. We already had flame wars over frameworks and this is way worse, your vibes vs. my vibes. Who would thought non-deterministic outputs would lead us here?
- sdesol 4mo ago> and no ways to measure whether AI is working better. What I do with my product is I explicity tell you to ask your agent. I have real world examples and real world repositories that you can try with: https://gitsense.com https://gitsense.com https://github.com/gitsense/smart-ripgrep https://github.com/gitsense/smart-ripgrep https://github.com/gitsense/smart-codex https://github.com/gitsense/smart-codex Token saving on average is not what I am mostly interested in though. I am more interested in knowing that the AI doesn't load unnecessary files in context, which can affect reasoning. You can just ask the agent after a task how many files do you think was not read by knowing the files purpose first?
- ziyasal 4mo ago> Mainstream CLIs and developer tools can easily ship a native --compact or --json-stream flag tailored for LLM consumption. Until they do, they won't soon , rtk, caveman, ponytail and many others are just trying to address every growing costs (for 2K org, its around 2.5M, for now), so these are trade-offs we are all know and adjusting, but unlike the author claims we know the trade-off well and forking these tools, benchmarking, verifying the output quality matches our needs and so on to make it work for us, so no blindly. For solo devs, yes, they might not really need it, self hosting another model to save would be better option. But for orgs thats a spicy part. Yes, its good that we see these articles are shedding some light but like we do with these tools, lets also consume these articles with a grain of salt.
- giancarlostoro 4mo agoI just typed in rtk gain on my Mac, unfortunately my main dev machine I reimaged due memory issues I had and it messing up a few things, but on my Mac I've shaved off roughly 51k input tokens, and 23k output tokens, and saved an average of 3 seconds per command. Not sure what the outrage is for or why they cared enough to write this up really. Not sure who is piping stacktraces through RTK, I only use it for very specific programs, shoving compiler output through it seems silly, but you can always instruct your agent to only use RTK for very specific sets of commands.
- cephei 4mo agoMany points about maintainability that this article makes seem to hold, especially with update and version output changes, but it doesn't even offer the simplest alternative. Most of these supported commands have flags to strip out noise and reduce output. Maybe agents aren't well trained on these. As a side note, has anyone tried a dual agent setup where the command output is proxied through a lightweight local model? I can imagine a scenario where all tool output is filtered through Qwen or similar locally to compact the tool output.
- jbellis 4mo agoI feel bad that I wasted my time reading this. On the points in the article: 1. Yes, "gain" is a vanity metric but it's harmless, nobody is being "fooled" here. 2. This could be a problem in principle, sure, but unless you're actually vetting bug reports you're just spreading FUD. 3. Again, do you have any reason to believe that the thousands of devs using rtk are silently tanking their performance without noticing? here's a thought: instead of reporting that SOMEONE SHOULD MEASURE THIS, you could, you know, measure it yourself. 4. Good lord, what is this doing in a purportedly technical article? 5. Yes, this is inherent in the problem domain, again, nobody is being "fooled". Yes, I'm grumpy; reading this article was a waste of time. Bias: had my first RTK pr accepted today, so I guess I probably know more about it than this guy who got offended by "gain" and spit out the first thoughts that came to mind.
- beepbooptheory 4mo agoHow is 1 not more damning? It sounds like the fundamental service they are purportedly providing is not real. Am I reading it wrong?
- lackoftactics 4mo ago1. Are you sure no one is fooled? It’s the main thing managers are praising rtk for and using as an argument for it’s validity. If this is gamed, then it paints a very different picture. 2. No, I didn’t vet all the reports. But they paint quite convincing picture of the problems present in the library, which has a very ambitious goals of handling every popular command and making it less verbose. 3. You know this is not a valid point. Engineers tanking performance and choosing based on hype is nothing new. Github stars and usage is not a valid argument, when the tool is not very transparent and could quietly fail. If it’s only couple percents less accuracy, most wouldn’t easily recognize it with the whole stack of skills, mcps and agents.md 4. Is it something more than a feature? If the benefit is $3 on $900 as other commenter pointed out using maybe better and well researched article than mine from codepointer, why would I risk that for all the possible bugs and worse accuracy. 5. Hard to address this one. Tough problem domain to handle with endless cli commands to capture and process properly. Congratulations on your accepted PR. I didn’t want to make you grumpy today. If you feel I am wrong, it’s very possible. I am just a guy who wrote my point of view, it doesn’t automatically make it valid. Once again sorry for making you grumpy.
- ilia-a 4mo agoFirst of all there is a way to made agents spot truncation by being aware of RTK compression and having bypass option (I use RTK_DISABLE=1) as a way of restoring original full text. Works fine, yeah it only compresses command output so only input tokens are affected in terms of "compression".
- RVuRnvbM2e 4mo agoAgree. I've watched agents go around in circles or use ridiculous workarounds after being confused by rtk output.
- akman 4mo agoAnybody have experience with https://github.com/chopratejas/headroom https://github.com/chopratejas/headroom? They seem to have similar goals in token reduction, but headroom appears to be broader in scope.
- Bnjoroge 4mo agoThis post offers virtually no data to back up their objections and reads as LLM-generated for the most part. I
- mumin00 4mo agoThere are ways to improve token usage but no tool will work right on all prompts everytime.
- stephantul 4mo agoWe’ve been on the receiving end of this complaint with Semble. I think it is a valid complaint, but constructing a benchmark for this kind of thing is just very difficult and expensive because of the (harness) x (model) x (mcp/cli) combination. With traditional ml/tooling, not showing benchmarks was usually a red flag. But for llm tooling, I’m not so sure.
- Otterly99 4mo agoI completely agree with this post. After I used it in one session of 300k tokens, I had maybe 3k tokens saved. Plus, if commits really are an issue for you in term of tolen consumption, you can always ask to hand over the reigns and apply the commits yourself as a rule (unless you're operating in a loop).
- wren6991 4mo agoYeah, RTK is problematic because of its focus on associations between kanji and arbitrary English keywords, many of which are poorly chosen and... oh it's an LLM thing.
- Zababa 4mo agoDon't worry, what is kept vs what is removed by the RTK LLM thing is just as arbitrary as the RTK English keywords with no measure of performance!
- nexroo 4mo ago[dead]
- citizenpaul 4mo agoIts interesting that you posted this now. I stopped using RTK about 2 weeks ago due to suspicions and some testing that it may actually be hurting my token usage due to it causing increased loops due to faulty responses to the LLM that confuse it. I only have some rough metrics, unfortunately speed of LLM work has derailed my attempt to nail down usage efficiency. I spend less than $200 per month on tokens anyway and my usage is not consistent. I won't really know for another 30 days when I look at the total billing per day. So far my token use has not increased. I also looked through the huge backlog of the RTK issues and got nervous.