9 ms·
> Opus 4.7 uses an updated tokenizer that improves how the model processes text. The tradeoff is that the same input can map to more tokens—roughly 1.0–1.35× de
by cupofjoakim 6mo ago
> Opus 4.7 uses an updated tokenizer that improves how the model processes text. The tradeoff is that the same input can map to more tokens—roughly 1.0–1.35× depending on the content type.
caveman[0] is becoming more relevant by the day. I already enjoy reading its output more than vanilla so suits me well.
[0] https://github.com/JuliusBrussee/caveman/tree/main https://github.com/JuliusBrussee/caveman/tree/main
- deleted 6mo ago[deleted]
- Tiberium 6mo agoI hope people realize that tools like caveman are mostly joke/prank projects - almost the entirety of the context spent is in file reads (for input) and reasoning (in output), you will barely save even 1% with such a tool, and might actually confuse the model more or have it reason for more tokens because it'll have to formulate its respone in the way that satisfies the requirements.
- make3 6mo agoI wonder if you can have it reason in caveman
- 0123456789ABCDE 6mo agowould you be surprised if this is what happens when you ask it to write like one? folks could have just asked for _austere reasoning notes_ instead of "write like you suffer from arrested development"
- Sohcahtoa82 6mo ago> "write like you suffer from arrested development" My first thought was that this would mean that my life is being narrated by Ron Howard.
- acedTrex 6mo agoYou really think the 33k people that starred a 40 line markdown file realize that?
- verdverm 6mo agoStars are more akin to bookmarks and likes these days, as opposed to a show of support or "I use this"
- zbrozek 6mo agoI use them like bookmarks.
- LPisGood 6mo agoI use them as likes
- giraffe_lady 6mo agoI intentionally throw some weird ones on there just in case anyone is actually ever checking them. Gotta keep interviewers guessing.
- andersa 6mo agoYou mean the 33k bots that created a nearly linear stars/day graph? There's a dip in the middle, but it was very blatant at the start (and now)
- pdntspa 6mo ago[flagged]
- embedding-shape 6mo ago> I hope people realize that tools like caveman are mostly joke/prank projects This seems to be a common thread in the LLM ecosystem; someone starts a project for shits and giggles, makes it public, most people get the joke, others think it's serious, author eventually tries to turn the joke project into a VC-funded business, some people are standing watching with the jaws open, the world moves on.
- simonw 6mo agoI was convinced https://github.com/memvid/memvid https://github.com/memvid/memvid was a joke until it turned out it wasn't.
- embedding-shape 6mo agoTo be fair, most of us looked at GPT1 and GPT2 as fun and unserious jokes, until it started putting together sentences that actually read like real text, I remember laughing with a group of friends about some early generated texts. Little did we know.
- Alifatisk 6mo agoAre there any public records I can see from GPT1 and GPT2 output and how it was marketed?
- walthamstow 6mo agoI don't think it was marketed as such, they were research projects. GPT-3 was the first to be sold via API
- deleted 6mo ago[deleted]
- embedding-shape 6mo agoHN submissions have a bunch of examples in them, but worth remembering they were released as "Look at this somewhat cool and potentially useful stuff" rather than what we see today, LLMs marketed as tools. https://news.ycombinator.com/item?id=21454273 https://news.ycombinator.com/item?id=21454273 / https://news.ycombinator.com/item?id=19830042 https://news.ycombinator.com/item?id=19830042 - OpenAI Releases Largest GPT-2 Text Generation Model HN search for GPT between 2018-2020, lots of results, lots of discussions: https://hn.algolia.com/?dateEnd=1577836800&dateRange=custom&dateStart=1514764800&page=0&prefix=false&query=GPT&sort=byPopularity&type=story https://hn.algolia.com/?dateEnd=1577836800&dateRange=custom&...
- egorfine 6mo agoThey are indeed impractical in agentic coding. However in deep research-like products you can have a pass with LLM to compress web page text into caveman speak, thus hugely compressing tokens.
- claytongulick 6mo agoI don't understand how this would work without a huge loss in resolution or "cognitive" ability. Prediction works based on the attention mechanism, and current humans don't speak like cavemen - so how could you expect a useful token chain from data that isn't trained on speech like that? I get the concept of transformers, but this isn't doing a 1:1 transform from english to french or whatever, you're fundamentally unable to represent certain concepts effectively in caveman etc... or am I missing something?
- egorfine 6mo agoGood catch actually. Okay maybe not exactly caveman dialect, but text compression using LLM is definitely possible to save on tokens in deep research.
- ieie3366 6mo agoAll LLMs also effectively work by ”larping” a role. You steer it towards larping a caveman and well.. let’s just say they weren’t known for their high iq
- DiogenesKynikos 6mo agoThis is why ancient Chinese scholar mode (also extremely terse) is better.
- Hikikomori 6mo agoModern humans were also cavemen.
- roughly 6mo agoFun fact: Neanderthals actually had larger brains than Homo Sapiens! Modern humans are thought to have outcompeted them by working better together in larger groups, but in terms of actual individual intelligence, Neanderthals may have had us beat. Similarly, humans have been undergoing a process of self-domestication over the last couple millenia that have resulted in physiological changes that include a smaller brain size - again, our advantage over our wilder forebearers remains that we're better in larger social groups than they were and are better at shared symbolic reasoning and synchronized activity, not necessarily that our brains are more capable. (No, none of this changes that if you make an LLM larp a caveman it's gonna act stupid, you're right about that.)
- adwn 6mo agoI thought we were way past the "bigger brain means more intelligence" stage of neuroscience?
- nomel 6mo agoAll data shows there's a moderate correlation.
- waffletower 6mo agoEven neuronal density is simplistic, and the dimension of size alone doesn't consider that.
- stingraycharles 6mo agoWhile the caveman stuff is obviously not serious, there is a lot of legit research in this area. Which means yes, you can actually influence this quite a bit. Read the paper “Compressed Chain of Thought” for example, it shows it’s really easy to make significant reductions in reasoning tokens without affecting output quality. There is not too much research into this (about 5 papers in total), but with that it’s possible to reduce output tokens by about 60%. Given that output is an incredibly significant part of the total costs, this is important. https://arxiv.org/abs/2412.13171 https://arxiv.org/abs/2412.13171
- ACCount37 6mo agoSome labs do it internally because RLVR is very token-expensive. But it degrades CoT readability even more than normal RL pressure does. It isn't free either - by default, models learn to offload some of their internal computation into the "filler" tokens. So reducing raw token count always cuts into reasoning capacity somewhat. Getting closer to "compute optimal" while reducing token use isn't an easy task.
- stingraycharles 6mo agoYeah the readability suffers, but as long as the actual output (ie the non-CoT part) stays unaffected it’s reasonably fine. I work on a few agentic open source tools and the interesting thing is that once I implemented these things, the overall feedback was a performance improvement rather than performance reduction, as the LLM would spend much less time on generating tokens. I didn’t implement it fully, just a few basic things like “reduce prose while thinking, don’t repeat your thoughts” etc would already yield massive improvements.
- AdamN 6mo agoYeah you could easily imagine stenography like inputs and outputs for rapid iteration loops. It's also true that in social media people already want faster-to-read snippets that drop grammar so the desire for density is already there for human authors/readers.
- altruios 6mo ago
- bensyverson 6mo agoExactly. The model is exquisitely sensitive to language. The idea that you would encourage it to think like a caveman to save a few tokens is hilarious but extremely counter-productive if you care about the quality of its reasoning.
- andai 6mo agoDoes this imply that if you train it on Gwern style output, the quality will improve?
- gwern 6mo agoUnfortunately, that is an oversimplification for a highly RLed/chatbot trained LLM like Claude-4.7-opus. It may have started life as a base model (where prompting it with correctly spelled prompts, or text from 'gwern', would - and did with davinci GPT-3! - improve quality), but that was eons ago. The chatbots are largely invariant to that kind of prompt trickery, and just try to do their best every time. This is why those meme tricks about tips or bribery or my-grandmother-will-die stop working.
- deleted 6mo ago[deleted]
- Waterluvian 6mo agoHelp me understand: I get that the file reading can be a lot. But I also expand the box to see its “reasoning” and there’s a ton of natural language going on there.
- reacharavindh 6mo agoThis specific form may be a joke, but token conscious work is becoming more and more relevant.. Look at https://github.com/AgusRdz/chop https://github.com/AgusRdz/chop And https://github.com/toon-format/toon https://github.com/toon-format/toon
- alex7o 6mo agoAlso https://github.com/rtk-ai/rtk https://github.com/rtk-ai/rtk but some people see that changing how commands output stuff can confuse some models
- micromacrofoot 6mo agoI mean we had a shoe company pivot to AI and raise their stock value by 300%, how can we even know anymore
- bombcar 6mo agoLemonade and blockchain rides again! Or was it ice tea?
- addandsubtract 6mo agoWe started out with oobabooga, so caveman is the next logical evolution on the road to AGI.
- causal 6mo agoOutput tokens are more expensive
- sidrag22 6mo agoI hesitated 100% when i saw caveman gaining steam, changing something like this absolutely changes the behaviour of the models responses, simply including like a "lmao" or something casual in any reply will change the tone entirely into a more relaxed style like ya whatever type mode. I think a lot of people echo my same criticism, I would assume that the major LLM providers are the actual winners of that repo getting popular as well, for the same reason you stated. > you will barely save even 1% with such a tool For the end user, this doesnt make a huge impact, in fact it potentially hurts if it means that you are getting less serious replies from the model itself. However as with any minor change across a ton of users, this is significant savings for the providers. I still think just keeping the model capable of easily finding what it needs without having to comb through a lot of files for no reason, is the best current method to save tokens. it takes some upfront tokens potentially if you are delegating that work to the agent to keep those navigation files up to date, but it pays dividends when future sessions your context window is smaller and only the proper portions of the project need to be loaded into that window.
- SEJeff 6mo agoI believe tools like graphify cut down the tokens in thinking dramatically. It makes a knowledge graph and dumps it into markdown that is honestly awesome. Then it has stubs that pretend to be some tools like grep that read from the knowledge graph first so it does less work. Easy to setup and use too. I like it. https://graphify.net/ https://graphify.net/
- sambellll 6mo agoSomeone should make an MCP that parses every non-code file before it hits claude to turn it into caveman talk
- xnx 6mo agoThere's a tremendous amount of superstition around LLMs. Remember when "prompt engineering" "best practices" were to say you were offering a tip or some other nonsense?
- OtomotO 6mo agoAnother supply chain attack waiting? Have you tried just adding an instruction to be terse? Don't get me wrong, I've tried out caveman as well, but these days I am wondering whether something as popular will be hijacked.
- pawelduda 6mo agoPeople are really trigger-happy when it comes to throwing magic tools on top of AI that claim to "fix" the weak parts (often placeboing themselves because anthropic just fixed some issue on their end). Then the next month 90% of this can be replaced with new batch of supply chain attack-friendly gimmicks Especially Reddit seems to be full of such coding voodoo
- xienze 6mo ago> coding voodoo Well, we've sacrificed the precision of actual programming languages for the ease of English prose interpreted by a non-deterministic black box that we can't reliably measure the outputs of. It's only natural that people are trying to determine the magical incantations required to get correct, consistent results.
- JohnMakin 6mo agoMy favorite to chuckle at are the prompt hack voodoo stuff, like, “tell it to be correct” or “say please” or “tell it someone will die if it doesnt do a good job,” often presented very seriously and with some fast cutting animations in a 30 second reel
- computomatic 6mo agoI was doing some experiments with removing top 100-1000 most common English words from my prompts. My hypothesis was that common words are effectively noise to agents. Based on the first few trials I attempted, there was no discernible difference in output. Would love to compare results with caveman. Caveat: I didn’t do enough testing to find the edge cases (eg, negation).
- ruairidhwm 6mo agoI literally just posted a blog on this. Some seemingly insignificant words are actually highly structural to the model. https://www.ruairidh.dev/blog/compressing-prompts-with-an-autoresearch-loop https://www.ruairidh.dev/blog/compressing-prompts-with-an-au...
- cheschire 6mo agoI suspect even typos have an impact on how the model functions. I wonder if there’s a pre-processor that runs to remove typos before processing. If not, that feels like a space that could be worked on more thoroughly.
- 0123456789ABCDE 6mo agothere is no pre-processor, i've had typos go through, with claude asking to make sure i meant one thing instead of the other
- PhilipRoman 6mo agoI strongly suspected that there was some pre/postprocessing going on when trying to get it to output rot13("uryyb, jbyeq"), but it's probably just due to massively biased token probabilities. Still, it creates some hilarious output, even when you clearly point out the error: Hmm, but wait — the original you gave was jbyeq not jbeyq: j→w, b→o, y→l, e→r, q→d = world So the final answer is still hello, world. You're right that I was misreading the input. The result stands.
- 6mo ago
- TIPSIO 6mo agoOh wow, I love this idea even if it's relatively insignificant in savings. I am finding my writing prompt style is naturally getting lazier, shorter, and more caveman just like this too. If I was honest, it has made writing emails harder. While messing around, I did a concept of this with HTML to preserve tokens, worked surprisingly well but was only an experiment. Something like: > <h1 class="bg-red-500 text-green-300"><span>Hello</span></h1> AI compressed to: > h1 c bgrd5 tg3 sp hello sp h1 Or something like that.
- Leynos 6mo agoCombine that with emmet / zen coding: https://en.wikipedia.org/wiki/Emmet_%28software%29?wprov=sfla1 https://en.wikipedia.org/wiki/Emmet_%28software%29?wprov=sfl...
- naoru 6mo agoYou'd like Emmet notation. Just look at the cheat sheet: https://docs.emmet.io/cheat-sheet/ https://docs.emmet.io/cheat-sheet/
- user34283 6mo agoI used Opus 4.7 for about 15 minutes on the auto effort setting. It nicely implemented two smallish features, and already consumed 100% of my session limit on the $20 plan. See you again in five hours.
- deleted 6mo ago[deleted]
- hayd 6mo agome feel that it needs some tweaking - it's a little annoyingly cute (and could be even terser).
- chrisweekly 6mo agoI really enjoy the party game "Neanderthal Poetry", in which you can only speak using monosyllabic words. I bet you would too.
- gghootch 6mo agoCaveman is fun, but the real tool you want to reduce token usage is headroom https://github.com/gglucass/headroom-desktop https://github.com/gglucass/headroom-desktop (mac app) https://github.com/chopratejas/headroom https://github.com/chopratejas/headroom (cli)
- kokakiwi 6mo agoHeadroom looks great for client-side trimming. If you want to tackle this at the infrastructure level, we built Edgee (https://www.edgee.ai https://www.edgee.ai) as an AI Gateway that handles context compression, caching, and token budgeting across requests, so you're not relying on each client to do the right thing. (I work at Edgee, so biased, but happy to answer questions.)
- gilles_oponono 6mo ago100% agree
- anandvshah 6mo agoI have used Edgee.AI and it is amazing.
- stavros 6mo agoI tried to use rtk for the same, and my agent session would just loop the same tool call over and over again. Does headroom work better?
- motoboi 6mo agoCaveman hurt model performance. If you need a dumber model with less token output, just use sonnet-4-6 or other non-reasoning model.
- hayd 6mo agoDoes it? I'm not sure I'd necessarily notice but I haven't found it noticeably worse.
- nickspag 6mo agoI find grep and common cli command spam to be the primary issue. I enjoy Rust Token Killer https://github.com/rtk-ai/rtk https://github.com/rtk-ai/rtk, and agents know how to get around it when it truncates too hard.
- ctoth 6mo ago1.35 times! For Input! For what kinds of tokens precisely? Programming? Unicode? If they seriously increased token usage by 35% for typical tasks this is gonna be rough.
- p_stuart82 6mo agocaveman stops being a style tool and starts being self-defense. once prompt comes in up to 1.35x fatter, they've basically moved visibility and control entirely into their black box.
- fzaninotto 6mo agoTo reduce token count on command outputs you can also use RTK [0] [0]: https://github.com/rtk-ai/rtk https://github.com/rtk-ai/rtk
- JustFinishedBSG 6mo agoInteresting, it doesn't seem intuitive at all to me. My (wrong?) understanding was that there was a positive correlation between how "good" a tokenizer is in terms of compression and the downstream model performance. Guess not.
- alach11 6mo agoOn my private internal oil and gas benchmark, I found a counterintuitive result. Opus 4.7 scores 80%, outperforming Opus 4.6 (64%) and GPT-5.4 (76%). But it's the cheapest of the three models by 2x. This is mainly driven by reduced reasoning token usage. It goes to show that "sticker price" per token is no longer adequate for comparing model cost.
- stacktraceyo 6mo agoWhat about some thing like https://github.com/rtk-ai/rtk https://github.com/rtk-ai/rtk
- willsmith72 6mo agoThat's such a poor way to communicate a number. I take it they mean an increase of up to 35%?
- 4b11b4 6mo agobut what about DDD
- ojuschugh1 5mo agotry this - https://github.com/ojuschugh1/sqz https://github.com/ojuschugh1/sqz