5 ms·
Transformers know more than they can tell: Learning the Collatz sequence
- niek_pas 10mo agoCan someone ELI5 this for a non-mathematician?
- poszlem 10mo ago[flagged]
- embedding-shape 10mo agoDo you think maybe OP would have asked a language model for the answer if they felt like they wanted a language model to give an answer? Or in your mind parent doesn't know about LLMs, and this is your way of introducing them to this completely new concept?
- NitpickLawyer 10mo agoFunny that the "human" answer above took 2 people to be "complete" (i.e. an initial answer, followed by a correction and expansion of concepts), while the LLM one had mostly the same explanation, but complete and in one answer.
- embedding-shape 10mo agoMaybe most of us here don't seek just whatever answer to whatever question, but the human connection part of it is important too, that we're speaking with real humans that have real experience with real situations. Otherwise I'd just be sitting chatting with ChatGPT all day instead of wast...spending all day on HN.
- pixl97 10mo agoIf life is a jobs program, why don't we dig ditches with spoons?
- NitpickLawyer 10mo agoOh, I agree. What I found funny is the gut reaction of many other readers that downvoted the message (it's greyed out for me at time of writing this comment). Especially given that the user clearly mentioned that it was LLM generated, while also being cheeky with the "transformer" pun, on a ... transformer topic.
- esafak 10mo agoThe model partially solves the problem but fails to learn the correct loop length: > An investigation of model errors (Section 5) reveals that, whereas large language models commonly “hallucinate” random solutions, our models fail in principled ways. In almost all cases, the models perform the correct calculations for the long Collatz step, but use the wrong loop lengths, by setting them to the longest loop lengths they have learned so far. The article is saying the model struggles to learn a particular integer function. https://en.wikipedia.org/wiki/Collatz_conjecture https://en.wikipedia.org/wiki/Collatz_conjecture
- spuz 10mo agoThat's a bit of an uncharitable summary. In bases 8, 12, 16, 24 and 32 their model achieved 99.7% accuracy. They would never expect it to achieve 100% accuracy. It would be like if you trained a model to predict whether or not a given number is prime. A model that was 100% accurate would defy mathematical knowledge but a model that was 99.7% would certainly be impressive. In this case, they prove that the model works by categorising inputs into a number of binary classes which just happen to be very good predictors for this otherwise random seeming sequence. I don't know whether or not some of these binary classes are new to mathematics but either way, their technique does show that transformer models can be helpful in uncovering mathematical patterns even in functions that are not continuous.
- jacquesm 10mo agoA pocket calculator that would give the right numbers 99.7% of the time would be fairly useless. The lack of determinism is a problem and there is nothing 'uncharitable' about that interpretation. It is definitely impressive, but it is fundamentally broken, because when you start making chains of things that are 99.7% correct you end up with garbage after very few iterations. That's precisely why digital computers won out over analog ones, the fact that they are deterministic.
- fkarg 10mo agoyeah it's only correct in 99.7% of all cases, but what if it's also 10'000 times faster? There's a bunch of scenarios where that combination provides a lot of value
- robot-wrangler 10mo agoI'll take a shot at it. Using collatz as the specific target for investigating the underlying concepts here seems like a big red-herring that's going to generate lots of confused takes. (I guess it was done partly to have access to tons of precomputed training data and partly to generate buzz. The title also seems kind of poorly chosen and/or misleading) Really the paper is about mechanistic interpretation and a few results that are maybe surprising. First, the input representation details (base) matters a lot. This is perhaps very disappointing if you liked the idea of "let the models work out the details, they see through the surface features to the very core of things". Second, learning was burst'y with discrete steps, not smooth improvement. This may or may not be surprising or disappointing.. it depends how well you think you can predict the stepping.
- rikimaru0345 10mo agoOk, I've read the paper and now I wonder, why did they stop at the most interesting part? They did all that work to figure out that learning "base conversion" is the difficult thing for transformers. Great! But then why not take that last remaining step to investigate why that specifically is hard for transformers? And how to modify the transformer architecture so that this becomes less hard / more natural / "intuitive" for the network to learn?
- embedding-shape 10mo agoWhy release one paper when you can release two? Easier to get citations if you spread your efforts, and if you're lucky, someone needs to reference both of them. A more serious answer might be that it was simply out of scope of what they set out to do, and they didn't want to fall for scope-creep, which is easier said than done.
- Y_Y 10mo agoFor interest, this popular pastime goes by several delicious names: https://en.wikipedia.org/wiki/Least_publishable_unit https://en.wikipedia.org/wiki/Least_publishable_unit
- kkylin 10mo ago:-) I don't question this decision is sometimes (often) driven by the need to increase publication count. (Which, in turn, happens because people find it esaier to count papers than read them.) But there is a counterpoint here, which is that if you write say a 50-pager (not super common but also not unusual in my area, applied math) and spread several interesting results throughout, odds are good many things in the middle will never see the light of day. Of course one can organize the paper in a way to try to mitigate the effects of this, but sometimes it is better and cleaner to break a long paper into shorter pieces that people can actually digest.
- Y_Y 10mo agoWell put. Nobody want salami slices, but nobody wants War and Peace, either (most of the time). Both are problems, even if papers are more often too short than too long.
- Onavo 10mo agoInteresting, what about the old proof that neural networks can't model arbitrary length sine waves?
- kirubakaran 10mo agoThat proof only applies to fixed architecture feed forward multilayer perceptrons with no recurrence, iirc. Transformers are not that.
- ChadNauseam 10mo agoI don't know that computers can model arbitrary length sine waves either. At least not in the sense of me being able to input any `x` and get `sin(x)` back out. All computers have finite memory, meaning they can only represent a finite number of numbers, so there is some number `x` above which they can't represent any number. Neural networks are more limited of course, because there's no way to expand their equivalent of memory, while it's easy to expand a computer's memory.
- Onavo 10mo agoHere's the paper for your interest https://arxiv.org/abs/2006.08195 https://arxiv.org/abs/2006.08195
- jebarker 10mo agoThis is an interesting paper and I like this kind of mechanistic interpretability work - but I cannot figure out how the paper title "Transformers know more than they can tell" relates to the actual content. In this case what is it that they know and can't tell?
- godelski 10mo agoI believe it's a reference to the paper "Language Models (Mostly) Know What They Know". There's definitely some link but I'd need to give this paper a good read and refresh on the other to see how strong. But I think your final sentence strengthens my suspicion https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221