5 ms·
Is the broken English an optimization or a byproduct of the model being developed in China?
by kurante 1mo ago
Is the broken English an optimization or a byproduct of the model being developed in China?
- minimaxir 1mo agoOptimization. Why use many word when few word do trick?
- stavros 1mo agoWhat I find funny about "why use many word when few word do trick?" is that it's only slightly shorter than the regular "why use many words when few words do the trick?"
- rapind 1mo agoI always figured that was part of the joke, because a writer came up with it, and a writer would know (I assume?).
- inopinatus 1mo agoThe latter is not a complete alternative, it is ambiguously conflating vocabulary scale with word count, and also, it is not as funny
- happycube 1mo agoless word better(, many words bad)
- anshorei 1mo agofew word better, many bad
- seunosewa 1mo agoIn my research it's a solid 10 - 20% shorter on average, free of charge. That's nice to stack on top of other optimisations.
- andsoitis 1mo ago> Why use many word when few word do trick? Be concise. OR Brief is best. OR Eschew verbosity etc.
- gjvc 1mo ago"Omit needless words." -- William Strunk Jr. and E.B. White., The Elements of Style
- gaigalas 1mo agoOptimization on a idiosyncrasy. The same thing that makes Claude repeat "That was the most important thing you said in this whole conversation" is what makes grug speak optimize on token usage. Real humans get non-primary information from word variation. It's reasonable to hypothesize that it has a role in thinking things, because it endures. Our languages need to breathe over time, and flourishing might be one of the aspects that allows that breathing space.
- TiredOfLife 1mo agoSee world
- acheong08 1mo agoWhen GPT-5.6-sol's reasoning traces were leaked, they also used "caveman speak". Definitely a token efficiency optimization
- beefsack 1mo agoI can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?
- Barbing 1mo ago"Neuralese"
- walrus01 1mo agosome people made a 'caveman' speak qwen as a joke https://huggingface.co/ProCreations/grug-27b https://huggingface.co/ProCreations/grug-27b
- gaigalas 1mo agoIt's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).
- walrus01 1mo agoPersonally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about. https://github.com/p-e-w/heretic https://github.com/p-e-w/heretic
- gaigalas 1mo agoI think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful. Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model. That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).
- ekianjo 1mo agoSaving tokens
- andsoitis 1mo agoMore intelligent and shorter: Maybe add a small cycling cap or helmet if it doesn’t obscure the head.
- walrus01 1mo agoqwen3.8-flash-next also 'thinks' like this in its thinking stage before output, watching it 'think' in opencode, but it produces syntax correct and grammatically correct code comments, changelogs and readme type files.
- AdamConwayIE 1mo agoLikely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too. Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more "meaning" into a short sequence than English often allows for. That's why "caveman" reasoning is a pretty good fit. There's a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you're artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it's trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself. It's not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be. There are few tests of this, but one example I thought was interesting was here: https://github.com/PastaPastaPasta/llm-chinese-english https://github.com/PastaPastaPasta/llm-chinese-english I wouldn't say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.
- armcat 1mo agoLess tokens. These models already overthink like crazy especially for complex tasks.
- MetroWind 1mo agoI sometimes just chat the LLM in Chinese for token efficiency.