3 ms·
I have a hunch (based on using the Kimi models to write some clojure) that the article's AST point is exactly wrong. I had to spend a lot of time cleaning up wh
by regularfry 2mo ago
I have a hunch (based on using the Kimi models to write some clojure) that the article's AST point is exactly wrong. I had to spend a lot of time cleaning up when it miscounted closing parens, which implies that while the LLM may be operating on the AST, mapping to and from the token stream is harder, not easier, when the individual tokens carry less information.
If the author is finding that it works well, I suspect there's something else (code or comment style, maybe) that's compensating for it which didn't seem worth mentioning.
- bryanrasmussen 2mo ago>mapping to and from the token stream is harder, not easier, when the individual tokens carry less information. this is exactly what I would expect. Also if you are training on code on the internet, what are the chances that you get these kinds of structural errors, especially on code in blogs etc.? Lisp is known for being easy to drop a paren on accident so you saying not closing parens jibes with what I expect, that the LLM would predict wrong every now and then about if it should put one in a particular place.
- malloryerik 2mo agoParen issues with Clojure probably mean your functions are too long? I like to keep mine down to 8-10 lines or less when possible, keep them flat and composable, use threading macro and then transducers for performance. At least that's how I read it. AST might not matter much either way, or in a stranger way, because the LLM's corpus and progression through code will give it a kind of shadow or grooves of an ast, but it isn't making or receiving any ast from this piece of code. Still, Elixir scores best on the TenCent AutoCodebench, by far actually, "despite" being like Clojure built with an AST and immutability. Clojure wasn't part of those tests but I use both daily with LLMs (mostly Codex) and imagine it's on par with Elixir. The repl is better than Elixir's. Both have serious strengths.
- 4xel 2mo ago> Paren issues with Clojure probably mean your functions are too long? I like to keep mine down to 8-10 lines or less when possible, keep them flat and composable, use threading macro and then transducers for performance. This is a separate issue. You give great advice for both humans and LLMs, and anything in between, but the fact that it misses parens at all demonstrates it is not reasoning at the AST level, at least not directly, and that's the point the person your replying to is making. An hypotetical NN trained to produce valid AST would most likely never get it wrong, it would likely even be given the whole stack of opened context as its input to generate the next token, not just the preceding text tokens, not unlike humans have with indentation and parens highliters. At this point it would be pretty hard to miss a paren.
- malloryerik 2mo agoOh I agree the LLM not reasoning at the AST level, and was trying to say I believed this even more strongly than the person I was replying to, but that it didn't matter if you coded or had the LLM code in an appropriate style for a lisp. And then I made a tried to hint at a further claim that the base LLM is not reasoning at all beyond its attention heads I think. As I understand it the corpus space itself -- meaning the relations between tokens and lexemes and so on -- contains the shape of what we call reasoning, so that the language itself + weighting , attention heads, is doing any "reasoning" at all unless the LLM directly starts a chain-of-reasoning where it talks to itself, and if it's doing that just for one's delimiters then one probably hasn't used the lisp very well. I was probably unclear and sounding like I thought the LLM was fundamentally a reasoning device. As far as I understand, "reasoning" or an internal model other than the the language (training corpus corpus) + weights only exists when an LLM does "self talk" either as sub turns, a strong but expensive hack, or as a result of multiple turns layering up context. My claim is that the model can get delimiters right despite not reasoning about them, but deeply nested. My sense is that the model doesn't need to reason to track until attention heads are overwhelmed by nested delimiters; does those sound right? Anyway super interesting conversation, and I do think I was giving less credit to LLM reasoning, as seems to me an LLM trained on AST might still get it wrong a lot. So I don't tend to think AST is something that in and of itself makes languages with ASTs any better. But... immutability, which is practical thanks to AST, is another story. And if I'm wrong about anything here please let me know; I'm not an expert!
- chriswarbo 2mo agoIndeed, I use LLMs on some hobby Racket programs, and for Emacs Lisp, and it always messes up parentheses; then burns tokens trying to count them over and over (feels like "the number or rs in strawberry" problem). I've found https://github.com/shcv/parenmedic https://github.com/shcv/parenmedic to be somewhat helpful, which diagnoses parentheses issues based on when they disagree with indentation, rather than simply counting. The fact this works indicates that LLMs are paying more attention to whitespace than "actual structure".