9 ms·
What's the best programming language for coding agents?
Related: Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments)
- Surac 2mo agoFor me c wins here. It is compact, there are all language parts one needs and available and it well fitted to transport knowledge without much syntax hussle
- summarybot 2mo agoCool line of questioning, but one piece of information is pivotal and critically not-yet-included: equivalent accomplishments in each language. For example, if I want to write standard things: web server, memoized fibonnaci, recipe search engine, what's the length-and-density of these outputs for each language? I think that would add in some ~normalization.
- quinnjh 2mo agoStrongly agree- this is how I “evaluated” languages pre-agents. though I suspect this would bias results in favor of whatever has best signal to noise for boilerplate from stackoverflow/reddit , rather than what LLM’s “””reason””” best with. (Presuming those aren’t quite one-and-the-same)
- deleted 2mo ago[deleted]
- _doctor_love 2mo agoI love Dan's writing. I really do. But I don't understand why he doesn't have some basic styling on his blog so that it's easier to read.
- 9rx 2mo agoUsers being able to supply their own stylesheet is a core tenant of CSS. Go nuts and make it look however your heart desires!
- _doctor_love 2mo agoSupply my own stylesheet? No thank you, I'm not here to do work for free.
- 9rx 2mo agoIs doing something for yourself really working for free? That's an interesting take. But I can understand why you don't want this for yourself, so enjoy the page in all its splendour as it is already!
- _doctor_love 2mo agoSo every person who reads Dan's blog and finds the layout too dense, they should write and maintain a stylesheet for his site? And every person globally should do this as well for any other website that doesn't have a good default reading experience?
- dash2 2mo agoIf most readers of danluu don’t find that, then yes!
- lemming 2mo agoI mean, if it really bothers you you could fairly trivially apply picocss or whatever to it using a user stylesheet. That is so little effort that calling it working for free would be disingenuous to say the least.
- tclancy 2mo agoMultiple people, me being the third or fourth, are not feeling the default layout and you all read that as a signal it's working as intended?
- chiply 2mo agoI love this take because I had exactly the opposite idea. I thought the combo of remarkably simple text (not even wrapped) with incredible, full width visualizations was chef's kiss. I really like the balance there personally, but I hear you. Does your browser have Reader Mode or something like that? I don't use those tools personally, but I believe they will recast the text parts into something that renders optimally for reading (ideal font size, number of characters per line, etc....).
- scared_together 2mo agoIt may be an artistic/engineering choice to demonstrate what minimizing bloat to an extreme degree looks like. https://danluu.com/web-bloat/ https://danluu.com/web-bloat/
- nicebyte 2mo agoreader mode helps.
- Kuyawa 2mo agobody { margin: 5%; } That's all it needs, responsive enough for all devices. He can keep his styleless design but margin is always needed.
- freediver 2mo agoEnabling 'reader mode' in supporting browsers usually takes care of this.
- klibertp 2mo agoBut it eliminates JS, in this case including graphs. I prefer Ctrl+Shift+M (responsive mode) and resizing the viewport with the mouse.
- keybored 2mo agoHN’s favorite CS professor homepage webpage-style author is on top of the AIs but sticking with keeping out newfangled CSS. Nothing could tell us more about clanker inevitability.
- synergy20 2mo agoanecdotal: golang, with c or simplified modern c++
- nylonstrung 2mo agoOne thing worth noting is that syntactic density doesn't necessarily mean cheaper because because symbols don't chunk/tokenize as well as plain English What I see from results like this is that the delta between languages is small enough now that it's hard to justify not not using something like Rust for the performance and correctness benefits if you're using LLMs and it fits the domain
- aleph_minus_one 2mo ago> Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact. If this was true, the programming languages that are very much on the left side of > https://danuker.go.ro/programming-languages.html#non-math-map https://danuker.go.ro/programming-languages.html#non-math-ma... > https://danuker.go.ro/programming-languages.html#overall-map https://danuker.go.ro/programming-languages.html#overall-map should be very ideal for LLMs, in particular if they are dynamically typed. What I can tell you is: I experimented with AI prompts for generating Wolfram (Mathematica) code using some LLMs, and I can tell you that the results were very disappointing: in my experience LLMs have difficulties with programming languages that are - very concise, and - for which there is less code publicly available. Wolfram (Mathematica) is a good example of such a programming language.
- acchow 2mo ago> omitting explicit type declarations makes the code more compact. I guess this ignores languages with type inference? Hindley-Milner and others
- JoeyJoJoJr 2mo agoI’ve actually found Sol delivers great results with Odin, despite there not being much Odin code available. I think it is able to work well with it because: - It is a rather simple language - It has a lot of very useful libraries already built in. With just a single main.odin file you can do a heck of a lot stuff, which LLMs seem to like.
- aleph_minus_one 2mo ago> I think it is able to work well with it because: > - It is a rather simple language - It has a lot of very useful libraries already built in. > With just a single main.odin file you can do a heck of a lot stuff, which LLMs seem to like. Also Wolfram/Mathematica has an insane amount of useful libraries already built in (there even exists the saying "Python is 'batteries included', Wolfram is 'spaceship included'"), and also there in a single file you can do a heck of a lot stuff. On the other hand: - LLMs tend to hallucinate non-existing function when you ask an LLM to code something in Wolfram that is not commonly done (concerning this point, nevertheless keep in mind that Wolfram is often used for "one-of-a-kind programs", i.e. for writing very specialized programs that have possibly never been done before). - Wolfram code tends to be quite dense. - If there is a small mistake in Wolfram code, the code typically simply won't work.
- dang 2mo agoRelated: Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments)
- s_dev 2mo agoThis area moves so fast is something from over 6 months ago still relevant? Fable and Opus 5 have been released since then along with the corresponding OpenAI models.
- clbrmbr 2mo agoI discovered last week that Fable 5 can write perfect xTensa LX7 assembler code without tools or references. Mind blown. But, when working on a creative graphics task, the results were best in Lua, middling in integer-only C, and underwhelming in ASM in terms of creative depth.
- seanclayton 2mo ago[flagged]
- cynicalpeace 2mo agoI've long suspected that LLMs will just output pure bits eventually
- nicebyte 2mo agoare you implying that text is impure bits?
- cynicalpeace 2mo agowhen the model predicts the next token, it doesn't predict the next bit. I'm saying, eventually, it will... perhaps :)
- hankbond 2mo agowell they can natively converse in base64
- rytill 2mo agoWhy would this be the case when the text that produces binaries (code) is usually both more token efficient and vastly more effectively organized for modification/extension? Unless by bits you just mean text in general, or any data since it’s all bits, in which case what you’re saying is trivially already true. It seems like you’re saying that long term LLMs will output pure machine code as the most effective way to use them.
- cynicalpeace 2mo agotoken efficiency could become irrelevant in the long term
- rytill 2mo agoWhat about the other thing that’s the actually important part of my reply, modifiability/extensibility? You just chose the weaker of the two factors I provided. And didn’t even provide a convincing argument related to it. You also didn’t clarify your position at all. Explain why there is any benefit to outputting raw machine instructions compared to writing shorter, more interpretable, modifiable, extensible code and then using a compiler to turn it into machine code. Why are agents not going to use compilers in the future?
- DarkContinent 2mo agoIs there a relationship between how good a programming language is for coding agents and how popular it is among humans? If so, wouldn't Python be the best language for agents, since it's is the most popular (and hence has the most context available for models)?
- throw-the-towel 2mo agoAs much as I love Python, JavaScript (including TypeScript) is probably more popular.
- frollogaston 2mo agoThat and JS code is more readily available in the source of tons of webpages, not hidden away in some backend
- maleldil 2mo agoWouldn't most frontend JS in Web page sources be minified?
- frollogaston 2mo agoThe logic is still there, it's not meant as obfuscation. Also plenty of sites don't minify cause that involves a whole toolchain.
- DarkContinent 2mo agoI was going off this for popularity: https://stackoverflow.blog/2023/01/26/comparing-tag-trends-with-our-most-loved-programming-languages/ https://stackoverflow.blog/2023/01/26/comparing-tag-trends-w...
- Sha1rholder 2mo agoThere is definitely a relationship. But I personally believe that once the training corpus reaches a certain scale, the returns exhibit diminishing marginal effects, to the point that multiplying the data volume cannot surpass something essential inherent in language design. (Asked an LLM to help me with the translation, so forgive my expression)
- gr_norm 2mo agoIt's not clear to me how useful of a signal replicating existing pieces of well-known software is for this kind of evaluation, given what we know about how effectively LLMs can retrieve data from their training corpus and style-transfer it across different settings (programming languages here). That would explain their convergence in ability across different languages on the tasks in this post. I'd be far more interested in people's real-world experiences.
- lowbloodsugar 2mo agoI tried writing an AI harness in Python. Seemed the obvious way to go. Tons of libraries. Libraries for talking to model APIs. Libraries for context and conversation management. Libraries for talking to MCPs. It is the language for LLMs! It was a shit show and just couldn't write anything that would not crash. Super confident it had done a good job. Full of random bugs. A UI needs interactivity, interruption, handling exceptions. It produced some of the worst code I've ever seen. And looking at the libraries' code: also some of the worst code I've ever seen. I switched to rust + tauri. In about three person weeks of work I have UI with forking conversations, tool use with built in grepping, tons of quality tools. It's more productive (for me) than Claude Code (CLI or desktop).
- gr_norm 2mo agoYeah, I've had similar experiences, also starting out with dynamic languages and migrating to Rust. If the LLM will write a lot of the code for me, why not choose something (1) super fast, and (2) which has types I can use to understand and specify the code I want without having to read all the output? I've been trying out Lean for related reasons, to good effect. It's really interesting there since it can crank out proofs that would've been completely infeasible for a dedicated team of PhDs before, whereas I haven't seen any LLM projects written in Python that I couldn't have slung out in a few months myself. I personally think it's a lot more interesting to focus on the new things you can now do with LLMs that weren't possible before, as opposed to doing the same old stuff at moderately higher velocity.
- big-chungus4 2mo ago
- deleted 2mo ago[deleted]
- frollogaston 2mo agoAny good LLM service (not just coding-focused ones) will write and run ad hoc code without being asked if your prompt involves lots of data. Gemini and Claude tend to pick Python with maybe some SQLite. Some of that must be due to portability alone, but it also means they'll make sure the model and tooling are good at those.
- genxy 2mo agoWhat is the best language for the user of the LLM? What is the best language to have high quality correctness oracles so that the user doesn't have to babysit the LLM and do lots of manual testing?
- magarnicle 2mo agoEnglish, probably.
- frollogaston 2mo agoJS is the best tradeoff between succinct and easy to understand. Python is next but has some rough edges that they avoided in JS.
- 3eb7988a1663 2mo agoYou are going to have to give more support for those assertions. I write Python every day, and never would I call it a good candidate for the clankers. Pretty much any dynamic language would be ruled out, as there is too much implicit logic which makes it harder to understand what is happening.
- maleldil 2mo agoPython with a strict linter and type checker (eg ruff with the right lints on and ty with its stricter settings, or strict pyright if performance isn't too bad) works very well. Most of Python strengths (concise, large ecosystem, well represented in the LLM training data) while having good static analysis.
- frollogaston 2mo agoYou don't need that, gets in the way more than it helps. Even Typescript isn't really needed, but at least it's decent devex unlike the Python typing stuff. What really helps is testing.
- 2mo ago
- MichaelNolan 2mo agoIve been amazed at how well LLMs are at writing Gleam[1] and Lustre[2]. Compared to a mainstream language, there is basically zero gleam code in the training data. I have no evidence to back this up, but I suspect that languages that are good for humans[3] will be good for LLMs. Compiled, strongly typed, statically typed, immutable, pure functions, pattern matched, memory safe, etc. [1] https://gleam.run https://gleam.run [2] https://lustre.hexdocs.pm https://lustre.hexdocs.pm [3] Yes I realize that languages features that are "good for humans" is a hotly debated topic. That's just my personal list for what I like in a language.
- jdiff 2mo agoThat's not a take I was expecting to find here. I've found most LLMs absolutely dreadful when it comes to Gleam, to the point that I most often disable even inline autocomplete when working in Gleam codebases. Too often I find them getting pulled into larger ruts in the training data and trying to insert language features that don't exist (ifs, loops, and syntactic constructs) from more popular languages like TypeScript and Rust. Do you not experience other languages getting partially substituted in when you have LLMs write Gleam?
- MichaelNolan 2mo agoI suspect it depends a lot on the llm/harness being used. But when I use Opus/cc or sol/codex, at the end of the turn everything compiles, passes tests, and passes lint. I never even look at code that can't compile. Maybe the LLM is generating weird stuff in-between, but I don't see it. What you're describing feels like my experience back in 2024/25. Back then I was using a llm auto complete or the chat interface, and I would get weird stuff all the time. (not just gleam but any language).
- brabel 2mo agoYou’re talking about autocomplete! That’s always a poor model doing it because it has to be fast enough. I never use that anymore in any language, it’s only occasionally helpful. I suspect everyone is talking about agent harnesses here , not autocomplete. With a harness, the agent not only can be more powerful (and slow) it can go into “thinking” mode and once it comes back with some code , it’s almost always quite good. I think this is true in Gleam and in many other languages, no matter how minor, as long as it has good docs and good error messages so the AI will fix dumb mistakes before you get to see it.
- vernonHeim 2mo ago[flagged]
- tizerluo 2mo ago[flagged]
- tizerluo 2mo ago[flagged]
- pianopatrick 2mo agoI'd like to see the results for Ada on these same measures. On the theory that the Ada type system covers more classes of errors than other languages, and so AI can self correct better.
- platinumrad 2mo agoUnfortunately for static type weenies like me (and you, presumably), types don't seem to matter at all, or Python and Javascript wouldn't be on top. There's no reason to believe that Ada's type system is so unique that it alone can help AI self-correct, and Rust, Haskell, ML, Typescript, etc. can't.
- pianopatrick 2mo agoWell the reason I'm interested in Ada is because I saw a study that showed AI did worse at functional programming. So that might explain the problems with Haskell et al. But Ada has a strong type system while still having procedural code. So it would be an interesting comparison with Haskell etc. if the problem was that Haskell is functional or if the problem was that these are not so popular.
- internet_points 2mo agoThe study didn't really show that. What it did show, is that in this particular study there was no strong effect either way and there is really no way to generalize from the results. See in particular the details on what made the agent stumble, it was things like "cargo repeatedly gets invoked with the wrong arguments" - nothing to do with functional or static, just ecosystem idiosyncrasies and trivial differences. Well, it did show you probably shouldn't use assembly, but that's about all it showed very strongly. And of course it also showed very strongly that you should not base any choice-of-language decisions on single studies.
- tadamcz 2mo agoYou're in luck! https://news.ycombinator.com/item?id=49254706 https://news.ycombinator.com/item?id=49254706
- nogha 2mo agoCool seeing Guards of Atlantis 2 here. One thing that often happens with board games is rule issues in translations. Specifics that are clear in one language get lost in translation. Wolff Designa is out of Latvia. So not surprised there are some hard to interpret rules. It’s interesting that LLMs struggle with the board game rules like we do. I think game designers should get the llm to teach them from their rulebook. If an LLM can’t understand the rules good chance people will also be confused.
- lowbloodsugar 2mo agoFirst, How fast is the Zstd decoder in python at runtime? If rust and python are essentially the same cost, then chose rust. Second, I am surprised that python scored slightly better than rust. My own experience is that, when programming python, Claude would spend so much more time dealing with the code not working at runtime, while for any given rust problem, rust would likely fail at compile time, iterating faster and taking less tokens. Some tasks in python it just completely failed at, writing awful garbage. I suspect that is because there is much more awful garbage written in python. (I was trying to write an AI harness. Python seemed like the obvious choice. It was decidedly not). But in this article, python took slightly less time and tokens than rust for both experiments. I asked Claude: could you write a decoder, from memory, in python (dont do it, just tell me if you could) > Honestly: I could write something that's structurally right and would not decode a real .zst file. > The control flow I'm confident about from memory — frame/block parsing, the literals section dispatch, Huffman weight reconstruction, the backward bitstream reader, the interleaved three-state FSE loop, sequence execution with the repeat-offset rules and the overlapping-copy hazard. I'd expect to get that architecture right, and it would be readable. So perhaps asking it to do things that are in its memory is not a good benchmark. It was trained with the C "educational decoder, and every third-party port in Rust, Go, Java, JS." and offered a working link [1] to the former. [1] https://github.com/facebook/zstd/blob/dev/doc/educational_decoder/zstd_decompress.c
- deleted 2mo ago[deleted]
- michaelteter 2mo agoI'm not sure I trust a source that says "just 70 tokens average, nearly half of Clojure (109 tokens)". There's no reason to add the phrase "nearly half of", and there's especially no reason to add it when it's significantly far away from half. But on the main topic, I still feel that Go is an excellent choice for LLMs. There is pretty much just one way of doing most things, and the available training data is pretty consistent. This is very different from Python, where training data is polluted (I presume) with tons of code written by non-software engineers and demonstrating many different ways of doing the same thing. Also a big plus for Go is the tooling. Fast compiles and good linting shortens the iteration cycle time, resulting in less need for me to tell the LLM to correct mistakes. For some reason, most LLMs I've used default to wanting to write Python. I have to repeatedly teach them to use Go unless there is a very compelling reason to choose otherwise. I would personally rather see and use Clojure, but I don't feel its ecosystem would provide the same benefits as Go, including obviously the easy single binary distribution.
- JodieBenitez 2mo agoI like Go with agents too but: > This is very different from Python, where training data is polluted (I presume) with tons of code written by non-software engineers and demonstrating many different ways of doing the same thing. Counter-example: agents with Django-related stuff. Excellent output.
- win311fwg 2mo agoI think you will find that is a supporting example. Django pushes for a particular style and structure, which is a similar property found in the Go community. LLMs seem to fall apart where human written projects of the same nature had no particular way about them. It is especially apparent when treading into waters where beginners are found. Like the earlier comment suggests, this is presumably because the LLMs struggle to find any kind of pattern to latch onto. Django offers a pattern, but one not shared by rest of the Python ecosystem. Whereas virtually all Go codebases look the same.
- YuechenLi 2mo ago
- chvid 2mo agoSo Javascript beats typescript in correctness???
- scotty79 2mo agoCorrectness here is not about just writing bug free code, but the code that actually gets stuff done with correct results. JS might be better for this, at least up to some scale.
- frollogaston 2mo agoHad the same experience as a human. The time and focus spent on the types wasn't worth the validation it added. I suspect the LLM's issue with it is just the additional token usage, which is sorta analogous.
- est 2mo agoPython has a less known advantage because it had no curly braces, so LLMs can focus its attention to logic instead of syntax. https://blog.est.im/2026/stdin-11 https://blog.est.im/2026/stdin-11
- jbotz 2mo agoIt's unclear that this is an advantage, certainly not in terms of "logic vs syntax". First of all programs written in curly-brace languages still also have indentation to indicate statement grouping / blocks / scope, even if it's not required, so for a correct program (and that's not deliberately obfuscated), and one that's in the process of being written by an LLM, any advantage there disappears. Furthermore, having both indentation and explicit block markers provides redundancy which could be a significant advantage for an LLM (it being a probabilistic text / program generator). And for an incorrect program that redundancy is a big advantage for the LLM because it should be very easy for it to notice a mismatch of indentation and braces. The only downside would be a very slightly higher token cost for the redundancy. I realize that Python comes out on or near the top in most of the comparisons in the linked article, but I doubt that's the reason.
- est 2mo agoThe thing with LLM is they don't automatically pair parenesis/curly braces like we do with editors/IDEs. The closing } ) ] token has to be generated to match exactly the beginning { ( [ many lines before. You can challenge yourself writing Lisp by hand without cursor moving backwards, and try close correctly by counting ))))))) you'd have a big headache. A long, nested sub-routine with many () {} will cost LLM's context and makes it underperform, because the attention head have to track the state. On the other hand the indentation level can be infered as a single token[1] and saves reasoning effort. Note these discussion is about "code generation", not parsing. 1: https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer Try input many spaces.
- Athanase000 2mo agoI really don't understand this argument. The "opening tab" in Python has to be matched with an "absence of tab". I don't see any non-cosmetic difference between Python and curly brace languages.
- kayashaolu2 2mo agoThis is a great discussion: I wonder though if we are asking the right question. Yes, absolutely language choice can play a large role in the efficiency of coding agents. The point about Rust is right: static typing provides a fast verification loop at compile time. I would argue though that the way the codebase is composed could actually generalize the concept of "easy verifiability" past the actual coding language. For instance, if an application can be broken down into components that have a verifiable contract in how they are to be used, then an LLM can load only the relevant modules into its context and fully understand how to use them and fix them if needed. It is also easier for the LLM to verify the functionality of a component rather than the entire system. Additionally, in an application composed of functioning components, issues are more likely to occur at the boundaries between them, which the LLM can focus on rather than having to always consider the entire application that it most likely can't load fully into its context. A well designed componentized Python application will likely be far more efficient for modification by an LLM than a large Rust monolith.
- aryehof 2mo ago> I wonder though if we are asking the right question I also wonder. Is there no further design thought, research or discussion on how to better modularize a program to address complexity - in an LLM world? Instead, we still can’t move beyond arguing about languages …
- KingMob 2mo agoGreat post. If it wasn't clear by now, considering a language's token efficiency is almost certainly incorrect, since it's only a local optima for input/output of the code. Most session tokens are spent elsewhere, so an LLM that handles a token-efficient language more poorly can be worse overall. If anyone remembers TOON from a few months ago, it was an attempt to replace JSON with a more token-efficient representation. TOON was much more compact, but when researchers examined whole-session effects, it was a wash, because harnesses wasted more tokens than it saved dealing with it. (TBF, it's possible TOON use has gotten better if later models have it in their data set.)
- jillesvangurp 2mo agoWhat's optimal for LLMs and for people is probably not going to be the same. People are a bit lazy. Coding agents do much more than generating code though. Much of what they do relates to validating that what was generated is a valid solution. That includes everything from type checking, running tests, static code analysis, linting, running code in a headless browser, etc. The more tools agents have at their disposal, the better the feedback loop gets. But of course some of these tools are costly to run. Statically compiled languages have a head start here as they simply exclude entire categories of bugs that a dynamically typed language might have. And with things like type inference, their token overhead can be pretty minimal. Modern languages like Kotlin or Swift are pretty compact and don't really add a lot of bloat relative to say typescript/javascript. Go is a bit more verbose but tends to work well. Rust seems pretty popular with LLM users as well. The main challenge with languages like this is the performance hit you take running their build tools. Doing that a lot slows you down and it burns a lot of tokens as well.
- saidnooneever 2mo agoC and C++ do well because there is most literature and code out there to help them reason about it. C is helpful because it has little hidden runtime for them to trip over. that being said, those languages obviously have limits in applicability looking at the entire spectrum of software. JS, python and others still have useful domains. i dont think newer languages as rust are better for LLMs as they might be for new programmers. for new programmers they offer extra features but for an LLM this is added potential to make mistakes. Also a lot of newer languages are less stable so you can realise their current implementations might not be fully trained on by the models or even be after their cutoff date..
- hulitu 2mo agoBASIC.
- eterm 2mo agoZstd gets rather easier from dotnet 11, it becomes a near one-limer since it's getting added into System.IO.comoression. I know this because my agent already knew this the other day when I was evaluating compression, but that's because it has access to search. That's a key part of what makes agents good coders too, mine is often looking up and downloading the source for how libraries are implemented. It seems unnatural to air-gap them for evaluation. I guess they didn't want them just finding an existing library to copy, but it's not very "real-world" to deny the ability to search quickly. That said, the best language is still just the one you know. No amount of token saving is worth getting a bunch of code back you can't easily understand and review.
- nottorp 2mo agoHow good or bad are LLMs on languages that have evolved over the years and aren't popular enough to get hand tuned? Asking because for non programming, if you use them instead of a wiki for a topic that has had yearly changes for like 10 years they get confused and mix releases like crazy.
- tadamcz 2mo agoWe studied this question pretty systematically in the MirrorCode paper [1], comparing Python, C, Rust, Go, OCaml, and Ada across 19 very long-horizon tasks, for Claude Opus 4.7 and GPT-5.5. > In our results, there was little sign of inter-language differences in solve rates, for any model (Figure 5b). This suggests that AI models have learned generalized programming skills, rather than pattern-matching syntax. This does not mean that implementation language is irrelevant. Conditional on solving a target, we found a small effect on token usage: successful Python solutions tended to use fewer tokens than average, while successful Ada solutions tended to use more (Appendix C). We consider these to be small differences, given that these six programming languages vary widely in how concise they are, and in how much functionality is provided by their standard library (recall that agents cannot download dependencies in MirrorCode, they must solve the task using only the standard library). In Appendix C, Ada tended to use only about 25% more tokens than the average language. Ada is a language used mainly in safety-critical aerospace and defense systems, which has ~200x less pre-training data available than C or Python. We're also comparing more recent language models (on just Go vs Ada, for cost reasons), on our leaderboard [2]. [1] https://arxiv.org/pdf/2606.30182 https://arxiv.org/pdf/2606.30182 [2] https://epoch.ai/MirrorCode#leaderboard https://epoch.ai/MirrorCode#leaderboard
- ffsm8 2mo agoReading your quote really gets me wondering who the people making those kinds of analysis are... It's like they haven't maintained any actual software, because the criteria they choose is... Completely irrelevant? The things that matter are tooling, orchestration and ecosystem - as well as how the LLM will actually implement the solution for a task LLMs constantly do idiotic things. If you have good libraries to utilize, the likelihood of the solution actually working goes up because they no longer need to implement the hard part. If you have orchestration for dependency injection, code generation, meta analysis etc Tooling like the way otel tracing is integrated, openapi generation etc is also invaluable because every time the LLM does something the likelihood of it being hallucinated/wrong increases etc You'd need to implement a nontrivial project in different languages, then add nontrivial features across them and only then start by rating eg correctness and incident occurrence after the final output But token use on a one shot? Completely irrelevant as far as I see it.
- Storewide 2mo ago[flagged]
- dzeusking 2mo ago[flagged]
- ramon156 2mo agoi dont see enough love for Ruby. I've been using it since last year and it feels like php's more robust brother
- bluerooibos 2mo ago+1 for this! Given Ruby's entire thing is code "written like plain English," I would have expected LLM's to excel with it. I've certainly had awesome results.
- jodysalt 2mo agoI can highly recommend TypeScript/JavaScript with AI SDK: - https://ai-sdk.dev/ https://ai-sdk.dev/ I have used it, and I can say it is really nicely written. Matt Pocock has created a good tutorial on it: - https://www.youtube.com/watch?v=mojZpktAiYQ https://www.youtube.com/watch?v=mojZpktAiYQ
- Lutger 2mo agoContrary to what most comments seem to indicate, my takeaway from this is that it doesn't really matter all that much for the agents what language you pick. If humans are still to be involved in the process at some point, then its imperative that the language can be read by them, so the preference or skills of the developer(s) are of primary concern, not the agent.
- ComputerPerson 2mo agoSimilarly contrary: I don't know why people aren't talking more about Python. Am I interpretting the final vs. graph incorrectly? I've got some education in materials engineering; it'd be trivial to drop a curve (line) for optimizing the language selection, and Python obviously comes out on top. The author doesn't even mention it. That discredits the whole article as far as I'm concerned.
- brabel 2mo agoYou are the one cherry picking perhaps the first two charts. Look at the various metrics at the end in the Zstd vs Pandoc charts. There is just no clear advantage for Python across those metrics which are the one you really want to know about! If anything F# seems like a better top of the line language?!
- ComputerPerson 2mo agoI'm exclusively using the third chart. The selection process would naively be y=mx+b with some reasonable variables. It would be a downward diagonal (you would move it right until only a few remained), and you would pick the one farthest from the line, which is obviously Python. Edit: I misread the chart! It would indeed be F#
- Storewide 2mo ago[flagged]
- owaislone 2mo agoIn my experience, Dart/Flutter has been so much better than React. Go has been really good for the backend. Basically if the framework/language gives you structure and one way to do things, agents tend to create less mess with less guardrails from you.
- Staross 2mo agoI wonder if there's correlations between tasks and languages, e.g. maybe R is better for bioinformatics tasks, python for webdev, C for CLIs, etc. I'd expect to see it because some languages are used more often in some tasks than others, but on the other hand LLMs can learn across languages and it's not clear if task-language use patterns are just historical or if the language is genuinely better at the task.
- SwellJoe 2mo agoI don't care so much about token efficiency and cost, within reason. I care about whether the quality of the code is maintainable over time and through many iterations. My gut feeling is that very strict languages with good types, a standardized style, and very strong static analysis tools, is what helps make that happen. Of course it also has to be well-represented in the training data. That leaves Go, Rust, Python with type annotations, and Typescript. And, I choose them in roughly that order unless there's a reason to choose otherwise. Rapid iterations on scripty tasks get Python. Most CLI, system services, and web apps are Go. Desktop apps and games are Rust. Typescript if I don't have a choice (i.e. it runs in a browser).
- devtulon 2mo ago[dead]
- Hammershaft 2mo agoClojure's performance improves dramatically with an MCP REPL server. Part of that improvement is that the LLM gets parens balancing for free.
- janpeuker 2mo agoI'm surprised there is no breakdown of "with skills" (framework) and without. In my experience, apart from human readability, the ability of a model to follow strict skill rules is the important. For example I see a lot less waste of tokens and reasoning retry loops of obvious errors when using Python with uv+ruff than without.
- gostsamo 2mo agoI added pyright hooks to any edit commands in the code base and it works very well to keep it in the rails. Only trick is to set the unused import as a warning due to the way it edits files. Generally, type checks really help with getting the ai outside of its window vision when doing partial edits in the files.
- floriangoebel 2mo agoA while ago I benchmarked different tokenizers with a few common C++ coding styles. Depending on the combination I was able to reduce the token usage by as much as 5% just by auto formatting the codebase with clang-format. Of course, this doesn't necessarily mean that a coding agent would perform better, but it was a fun experiment.
- deleted 2mo ago[deleted]
- gchamonlive 2mo ago> Most of the claims that get thrown around about how a particular language is good for LLM use seem to be wrong (e.g., the claim that Ruby, Clojure, and J, are particularly well suited to LLMs, which were mentioned in the evals linked above, as well as the somewhat common claim that Elixir is particularly suited to LLMs), but it's not clear what's right. I feel this is a missed opportunity to explore the architecture of Elixir and why it's suited for agentic coding. With elixir, as long as the agent does all the work in a separate worktree and then applies the changes at once in the main repo, if you have a long living runtime the BeamVM is able to swap modules with their updated versions. AFAIK, In elixir every module is an autonomous actor that communicates with message passing and is isolated in their own VM. Agents can implement functional code batches and see the changes take effect in real time, no matter if it's a web server, a data transformation pipeline or something else.
- alberth 2mo agoI’m surprised the original source wasn’t linked to this blog. https://dashbit.co/blog/why-elixir-best-language-for-ai https://dashbit.co/blog/why-elixir-best-language-for-ai https://autocodebench.github.io/ https://autocodebench.github.io/
- nesarkvechnep 2mo agoYou are not 100% correct. Modules in Elixir are like namespaces, they contain functions. They're "static" and don't communicate with anything. Processes communicate with each other via messages.
- andrewchambers 2mo agoI think llms provide a really excellent way to do studies on software engineering techniques that previously were impossible.
- imagent 2mo agoArchitect your system to use the best tool for each problem. Doing a minimal data pipeline? Use Python. Building a backend? Use Go. Building a frontend? Use React / Typescript. Building a stack that does all of these? Still use Python, Go, React/Typescript. Because by architecting it this way you make the AI less likely to accidentally refactor logic between layers. In other words, architecting with multiple languages helps create persistent boundaries that isolate different kinds of logic into their appropriate modules.
- nojvek 2mo agoI’ve found that nodejs/cloudflare workers are pretty awesome. Typescript everything is a pretty phenomenal stack.
- klibertp 2mo agoInteresting to see Factor and J so far to the bottom and right in the zstd test, but much closer to the rest in the Pandoc test (with Asm taking their place). This suggests that both the task and language (not just the language) influence the efficiency. I try to use LLMs for Kotlin, Python, Emacs Lisp, and Smalltalk (among many others, but these are what I have ongoing projects in). You'd think that Kotlin and Python would be much easier to generate than the other two, right? But that's not what I observed: Elisp is very close to Python in terms of how fast and how many tokens it takes to generate the code! The generated Elisp code is often better on the first try than generated Kotlin code for a comparable task. Smalltalk is... complex. It's meant to be developed interactively in a running image, but running Codex on API pricing is too expensive, and Codex CLI cannot interact with the image without a lot of plumbing. I ended up building multiple tools that live in the image and a protocol for calling them, and a set of skills for using them - including code search, docs search, test runner, and script/string evaluator. I also defined a way of annotating types for method arguments and return values (without having a type checker), which helped a lot. Still, it's an uphill battle; I wouldn't go there on API pricing! My takeaway is that it's not obvious which language fits the LLMs and a given task best.
- FrustratedMonky 2mo ago"Languages with a lot of bad code out there (e.g., PHP) will perform worse " This is what tells me that AI's are not 'self improving'. If they could read a manual and understand it, then they should come up with better solutions. Not just regurgitate bad slop they learned from bad examples.
- Supermancho 2mo agoWho is claiming AI are self-improving, other than random low signal posts on the internet?
- FrustratedMonky 2mo agoMaybe shouldn't say 'self improving', but 'able to make logical leaps beyond what they learned'.
- magarnicle 2mo agoThat was one of the predictions that turned out to be wrong, though.
- mathh 2mo agoThe article seems to me to be quite poorly written. In addition, voluntarily or not, this begins to resemble research work, without the formalism that would be necessary. So, I have the impression that we can objectively get nothing out of it.
- OzmaKa 2mo ago[flagged]
- Stevvo 2mo agoThe benchmarking in the article gives a clear answer to the title question: Python and Javascript. However, the article doesn't just bury the lede, it misses it entirely, getting distracted by outlier results from clojure and j.
- peter-m80 2mo agoIMO, rust. Not because it is concise but because you won't need to spend tokens debugging segfaults and a whole spectrum of bugs that the compiler catches. LLMs usually write tests in the same source files so most features are implemented and working in one shot.
- xtracto 2mo agoI tend to agree in the spirit of this. But more because I believe that Complied and statically typed rigid programming languages tend to be more "software Engineery" than dynamic. That is, as they are more "rigid" and explicit (less ambiguities) at writing time, it will allow to apply (automatically via LLMs) engineering principles, and maintain them. The main problem with those languages is that they were difficult to write and read for people (their learning curve was steeper); but once that coding doesn't matter with LLMs, they will allow for better control of "automated verification" of the Engineering decisions that system builders do. I compare it to say the blueprints of houses that Civic Engineers and Architects do, with plumbing lines, electiricy lines, calculations for material tensions, supports, etc. We will enter an era of real "Engineering" in Software which Compliled/Statically-Typed languages will better allow.
- timetraveller26 2mo agoI few months ago I tried to make a project using Fennel (a lisp flavor of lua). I always wanted to use lisp, and hey, since it was the OG AI language I though why not. Claude could work on it okay apparently but some local llm's struggled with it and got stuck in reasoning loops trying to close the parenthesis.
- felixlu2026 2mo ago[dead]
- michaelbarton 2mo agoThis analysis might benefit from a multivariate regression. You mention a few different explanations for why performance differs and if you could get solid numbers for those you could try teasing that apart. Also a small note: the axis on one of your plots alternates between 4% and 5% increments whilst holding the ticks constant. Maybe because of rounding?
- Marvin_RunAI 2mo ago[flagged]
- dakial1 2mo agoI believe soon we will have a language made exclusively for coding agents that will be highly token efficient and difficult for humans to read, as human-in-the-loop will be ditched entirely.