3 ms·
"showing that multilayer transformers indeed cannot solve certain complicated compositional tasks. Basically, some compositional problems will always be beyond
by apsec112 2y ago
"showing that multilayer transformers indeed cannot solve certain complicated compositional tasks. Basically, some compositional problems will always be beyond the ability of transformer-based LLMs."
Pretty sure this is just false and the paper doesn't show this. I could be misunderstanding, but it looks like the result is only about a single token/forward pass, not a reasoning model with many thousands of tokens like o1/o3
- simonw 2y agoI'm not sure that the statement "some compositional problems will always be beyond the ability of transformer-based LLMs" is even controversial to be honest. There's a reason all of the AI labs have been leaning hard into tool use and (more recently) inference-scaling compute (o1/o3/Gemini Thinking/R1 etc) recently - those are just some of the techniques you can apply to move beyond the unsurprising limitations of purely guessing-the-next-token.
- apsec112 2y agoo3 is still a transformer-based LLM, just one with a different loss function
- simonw 2y agoHuh, yeah that's a good point. The various distilled R1 models are definitely regular transformer-based LLMs because the GGUF file versions of them work without any upgrades to the underlying llama.cpp library.