5 ms·
LLM don't work on 'compressed data'. LLM compress data into their latent space which allows them to become general. They learn the concept of things and how to
by Yopolo 2mo ago
LLM don't work on 'compressed data'. LLM compress data into their latent space which allows them to become general.
They learn the concept of things and how to do them because this is better compression than learning concepts one by one.
Which means, if an LLM 'learns' the concept of a poem, it can put everything into the formad of a poem instead of learning a billion poems.
- discreteevent 2mo ago> They learn the concept of things and how to do them because this is better compression than learning concepts one by one. When anthropic looked at how an LLM does addition it found it had some mental math heuristics that might or might not always work. The LLM hadn't learned the concept of addition. It had learned some heuristics that might work for some numbers. The result is that LLM's cannot add numbers reliably because they have not learned the concept of addition.
- Yopolo 2mo agoIt learned a concept of a heuristic which made it smart enough for the learning reward. Might be an architecture issue or a parameter size issue that it didn't learn to do math like a caculator. But look at your own math skills: How many numbers / how big of numbers can you keep in your head? How far is this heuristic away from how much a human learned until you start using pen and paper or a caculator?