4 ms·
Given that the most important feature for agent performance is the popularity of the language, ie, the amount of training data, (https://danluu.com/pl-tokens/ h
by a2ff6eeb0 2mo ago
Given that the most important feature for agent performance is the popularity of the language, ie, the amount of training data, (https://danluu.com/pl-tokens/ https://danluu.com/pl-tokens/), why would you cause problems for yourself by using Lisp rather than Python/Javascript if you care mainly about results fast, or C/C++/Rust if you care about performance too?
- jbott 2mo ago> the most important feature for agent performance is the popularity of the language This is explicitly called out as only weakly supported in that blog post: - You should use a popular language - There's weak support for this statement
- a2ff6eeb0 2mo agoFeel free to do your own analysis -- my informal experiments backs this up, though. I see worse results when I try to do anything in an unpopular language. It makes sense, needing to train the model on things that aren't already in its weighs takes up valuable context. Until we have models that update their weights based on what they've seen in their recent sessions and learn like people, this will be a problem. For now, though, between the results I'm seeing here, and the lack of need to look at code, I think this kills off any reason for me to use less popular languages.
- jauntywundrkind 2mo agoYour attempt was probably pretty, ahem, weak. How much support would you say you gave your goes, before you threw in the towel? I don't think there's really ever a downside to leaning in and making use of a language or system that works for you. Trying to tell people they should just use the popular thing is, imo, bad advice to turn hackers and experimenters into boring people.
- a2ff6eeb0 2mo agoA day or so for each of the oddball languages; again, I'm still waiting for an argument on why there's any value here, since the entire point of an agentic system like this is that I don't have to read the code. Experiment with the AI, sure, but you've got a pretty high burden of proof to show that AI is going to pick it up without a high per-prompt token cost. AI changes the constraints here for now, since it can't permanently learn things. I'm waiting until that changes, but right now it's better to use what it knows out of the box if you want good results. A better language doesn't buy me anything other than performance; the reason to stick an AI in here is to remove interactions with the code. I don't care what the AI chooses to use, as long as it gets results.
- wild_egg 2mo agoIn this case, the better language buys you increased iteration speed in addition to performance, and that is worth a lot.
- a2ff6eeb0 2mo agoWhy? I'm giving the system the same prompts either way.
- rdb_ 2mo agois it that llms write "better" typescript than let's say elixir because it has seen more of it..? or is it that you're relying on something like effect-ts to keep llms from tripping over even small things? coincidentally, "good code" in popular lang is rarely directly attributed to only that part; and it's also about the underlying principles it tries to follow in the code... another example; is it typescript that's good, or are "types" inherently making things/feedback loops easier to reason about in llms? (only using ts here for all example because it's probably one of the most "trained on" pl)
- a2ff6eeb0 2mo ago
- magnusi 2mo agoThere is a Common Lisp pro you are not seeing and that is that it has by far the best OOB debuggability/introspectability (especially when using SBCL) out of any practical language, while still having great performance
- khalic 2mo agoThere are many other confounding factors here, the type of prompting, how familiar you are with the language idioms, the context you gave, random bad quality runs, etc. You can’t tell that with a few uncontrolled runs
- evanjrowley 2mo agoI see you're interested in avoiding the need to read any code. It might surprise you to learn that autolith is very capable at reading and updating it's own code. The captured sessions at the linked page are three examples of this.
- magnusi 2mo agoCorrect, that's why I made it! :)
- yogthos 2mo agoThe obvious reason is that Lisp is perfectly suited for writing self modifying programs in a way pretty much no other language is. And as others pointed out, the evidence that agents work better with other languages is pretty thin. I've been using Claude, GLM, and DeepSeek with Clojure for around a year now, and they certainly do just fine in my experience. In fact, I've had much easier time maintaining LLM assisted programs in Scheme and Clojure than other languages I've tried using because functional style naturally leads to low coupling. And that makes controlling context far easier than the rats nest of shared state that you have in imperative languages.
- wild_egg 2mo agoLLMs have been solid at writing Common Lisp since Sonnet 3.5 and have been near flawless since the Opus 4.5 release. The niche language thing is really not a problem at all any more. If you're working in some esolang it doesn't take more than a 1-2k token primer in the context to get great results, and lisp is popular enough to not even need that. The benefit of having the agent directly in the image like with Autolith here is that it can directly inspect all defined symbols and explore and orient itself automatically. Really doesn't need much guidance to get great results.
- magnusi 2mo ago(author of autolith here) This all correct, I'd also add that in my experience, the GPTs are even better at Lisp, namely in the counting parentheses department. Which is not an issue that much per-se because in Autolith, the harness detects Lisp file edits (CL, Scheme, Clojure) and gives hints when the edits lead to unbalanced files (The heuristic is pretty simple, we detect if there's a mismatch, and if yes, it provide hints where the extra/missing might be based on indentation)
- a2ff6eeb0 2mo agoBut LLMs already do that with text, don't they? And I don't really want to interact with the code directly, so I'm not sure why I should care what language is used other than raw performance and LLMs ability to use it. Do you have benchmarks on non-trivial tasks (say, generating zstd) that show it does any better than rust?
- armitron 2mo agoThis is not just false but egregiously wrong. The regular syntax of Lisp is a tremendous asset when it comes to LLMs being able to work directly in the image. If I had to score languages by how well they work with current LLMs, Common Lisp and Emacs Lisp would be at the very top.
- magnusi 2mo agoIt's not just the syntax, but also the tools for debugging, the image paradigm itself, and the iterative approach to development! Autolith can spawn managed Lisp REPLs either from saved images (so it can do checkpoints) and triage changes before committing them to files, and then run test suites in the same REPL, it's been very useful for this.
- a2ff6eeb0 2mo agoDo you have any benchmarks for larger tasks? The best others here have claimed is that it's not strongly proven to be harmful when you look at benchmarks.
- magnusi 2mo agoWell, I made Autolith in Common Lisp because I like Lisp and I think it's by far the best and most practical language for self-modifiable live image agents. I have been trying Scheme and CL with LLMs for the last three years or so, and in recent months, I have finally decided that they are good enough. My idea is that well, it's good enough that I can now produce more training data just by using Autolith with the most basic claude/gpt subs, haha