3 ms·
I'm shocked to see how poorly these models, which I find useful day to day, do in solving virtually any of the problems in Unlambda. Before looking at the resu
by bwestergard 7mo ago
I'm shocked to see how poorly these models, which I find useful day to day, do in solving virtually any of the problems in Unlambda.
Before looking at the results my guess was that scores would be higher for Unlambda than any of the others, because humans that learn Scheme don't find it all that hard to learn about the lambda calculus and combinatory logic.
But the model that did the best, Qwen-235B, got virtually every problem wrong.
- __alexs 7mo agoThey are also weirdly bad at Brainfuck which is basically just a subset of C.
- astrange 7mo agoBF involves a lot of repeated symbols, which is hard for tokenized models. Same problem as r's in strawberry.
- bwestergard 7mo agoInteresting. So why do the models seem to handle deeply nested Lisp expressions just fine?
- kgeist 7mo agoProbably because there's a ton of code that deals with nested parentheses across languages in the training data, and models have learned how to work around tokenization limitations, when it comes to parentheses.
- astrange 7mo agoIt's because the models wouldn't work for coding if they couldn't do nested scopes, so people don't release models unless they work. They can only do it in a limited form though, because transformer models only have limited "memory". I don't think they can fully implement parsing.
- culi 7mo agoYeah well they also still struggle with "4 + 6 / 9" so I'm not sure why anyone is surprised with these findings
- joshmoody24 7mo agoThis surprises me too. I've experimented with using LLMs to convert lambda calculus expressions into combinatory logic. There is a simple deterministic way to do this, and LLMs claim to know it, and then they confidently fail.