4 ms·
So the interesting thing is that this shrinking “linguistic diversity” is fundamental to how an LLM works. The LLM is a big probabilistic statistical trick. It
by ohyes 25d ago
So the interesting thing is that this shrinking “linguistic diversity” is fundamental to how an LLM works.
The LLM is a big probabilistic statistical trick. It picks the next token based on certain words are simply “the best” because they are specific and well connected to other tokens. The is gives them a great overall cost function. (Basically a good score on “will it make sense in context” while also having specific meaning that makes it better than other options, unambiguous in common use and being a single token rather than several).
You can trim those tokens, but then you just get other tokens that are “the best” tokens (and you’re worse off because the output became less clear).
The cool thing is this seems to get worse the more powerful and accurate your model is, because it is picking technically / statistically perfect tokens, not tasteful ones.
- DoctorOetker 25d agothat's a very longwinded claim that inference providers are sampling with temperature T=0, but is that even true? a sufficient explanation would be merely sampling at a lower temperature compared to human sources providing similar content
- lensecat 25d agoIt's still probabilistic from a huge dataset while the human mind is not.
- BoredomIsFun 25d agoIt does not matter really - stiffness does lower up to T=0.7, then platoes; even at high temperatures tics/slop-patterns are still there.
- orbital-decay 25d agoThat's pretty model-specific, for example DeepSeek of the v3/R1 era would already start losing coherence occasionally at t=0.7 with no other samplers
- orbital-decay 25d agoIt's not fundamental at all, that's what randomized sampling is for. Try tinkering with a base model and you'll be surprised how diverse it is. The semantic collapse happens in post-training that is using the current methods.
- BoredomIsFun 25d ago> Try tinkering with a base model and you'll be surprised how diverse it is. Have you tried? I have. Not much different from RLHFed; full of tics and slop, similar but slightly different from intsruction posttrains.
- orbital-decay 25d agoYes, although admittedly I haven't tried recent ones which have a bunch of synthetic data in them (mainly to aid the reasoning), and are usually only available after mid-training. One look at the logits/output distribution and it's clear the base is pretty different.