4 ms·
Something like this seems expected, right? If you tune a statistical model to very high accuracy in "addition" over tokens, then the resulting structure of the
by lsy 2y ago
Something like this seems expected, right? If you tune a statistical model to very high accuracy in "addition" over tokens, then the resulting structure of the model must correspond to some structure in the training data. And fourier would make sense for some token like "123" which internally is represented as the integer 7633, but needs to "contain" information about the text digits for math to work. Notably, this still ends up being in some way a statistical endeavor rather than truly learning addition, as even the fine-tuned model doesn't reach 100% accuracy.
- godelski 2y agoThere's always a lot of research that is "expected", but there's nothing wrong with that. The two most common reasons this happens are: 1. Well somebody has got to do the work and we can't all just go around assuming stuff, even if we're pretty confident. The confirmation helps and is beneficial to the community. 2. It's obvious post hoc and you have gaslit yourself into thinking that you already knew it because you kinda knew it at a high level and you only read the result at a high level too so you entirely miss all the actual details and all the context (especially since there is so much context that never makes it into a paper[0]) 3. (bonus) It addresses the same thing that someone else addressed but from a different approach and the new approach can provide additional insight. Either way, it is beneficial to the community. Sure, nothing is groundbreaking but that's how science is. 99% incremental steps. And hey, these days most ML papers are just an active demonstration of how well you can brute force search optimal hyperparameters (i.e. how much compute you can afford). I see far fewer sufficiently isolating variables and actually provide strong evidence of the things claimed as novel, not recognizing that benchmark results are far from sufficient. But I blame reviewers for that, but also see rant in [0] [0] I think the hardest thing about beginning a PhD is that you're reading a bunch of papers going like "why the fuck are they doing this?" and the problem is that you don't have enough breadth or depth to get it. You don't understand the decades long conversation of how we got here and what problems were being addressed along the way. To be fair, a lot of this is never stated explicitly and so you annoyingly have to piece everything together by reading a few hundred papers. But also, good luck providing all that context within page limits and besides, papers are written for /peers/, by which I mean niche peers, not domain peers. Ain't nobody got time to write textbooks, because you're just trying to publish so you don't perish and you're already exhausted from all the grant writing, rebuttals, bureaucratic work, and all that fun jazz.
- deleted 2y ago[deleted]
- wongarsu 2y ago> Notably, this still ends up being in some way a statistical endeavor rather than truly learning addition, as even the fine-tuned model doesn't reach 100% accuracy If that's our metric then most humans haven't truly learned addition either For any neural network, the standard you can expect for any learned skill is closer to a human learning that skill than to a computer programmed to do that thing. There will be occasional mistakes
- Bolwin 2y agoWell we have formally learned addition but most of the time I actually do it, I'm not doing it, I'm going based on some half remembered pattern checked with statistical expectations. I'm sure the llm could formally do it too
- mannykannot 2y agoHow LLMs tackle addition is an interesting question in its own right, independently of whether their accuracy provides a metric for judging their ability relative to that of typical humans.
- bongodongobob 2y agoSeems closer to "truly learning addition" (whatever that means) than what humans do. We use a mechanical algorithm to carry the 1 etc.
- pinkmuffinere 2y agoWhat? This is a crazy take. I feel we (humanity) has truly understood addition to a very high degree. The fact that we can see what an LLM is doing and understand “oh it’s doing addition in a rather roundabout way via Fourier transforms” is a testament to just how well we understand addition — we can recognize hundreds of different ways to achieve the operation, know that they are equivalent, and pick the most convenient one for the situation
- Ukv 2y agoI think humans do two broad types of arithmetic: #1: For small numbers, the answer is "directly" available to us without consciously applying any steps #2: For larger numbers, we consciously apply some formal method to break the problem down into smaller steps of type #1. Like column addition, adding digit-by-digit and carrying the ones Despite seeming direct, I'd argue that if we could see at a low-level what our neurons were actually doing for #1, like to get 3 + 5, it would likely also look roundabout. Might even be a similar process as with LLMs, approximating magnitude (~7-9) then snapping to parity (even, since 3 and 5 are odd). LLMs should be capable of #2, including choosing an appropriate method, with chain-of-thought reasoning. But in addition to that, and I think what bongodongobob is getting at, is that LLMs appear to have a more robust #1 than us - being able to accurately add far larger numbers whereas we'd normally fall back to a step-by-step method after one or two digits.