15 ms·
Read the first few comments and surprised I didn’t see it, but training data. The voluminous amount of Python in the training data. I could write in brainfuck
by _boffin_ 5mo ago
Read the first few comments and surprised I didn’t see it, but training data. The voluminous amount of Python in the training data.
I could write in brainfuck with ai, but I presume, wouldn’t get the same results than if going with python.
My follow up question: with AI now, why care about a lang until you need to?
- jryio 5mo agoI wrote about the meta thesis of programming languages in the training data here https://jry.io/writing/use-boring-languages-with-llms/ https://jry.io/writing/use-boring-languages-with-llms/
- _boffin_ 5mo agoPlease distill instead of having me navigate off site. Include link for additional info. edit: side -> site
- btown 5mo agoAlso, every single interpreter error has an entire corpus of StackOverflow-esque fix suggestions alongside it, and the model has been fine-tuned to minimize such errors on the first try. This hasn't been done for more obscure languages. You'll likely take more turns, on average, to get a working output, even if your problem is fully verifiable via test input/outputs - and if it's not verifiable, you don't want the "attention" of the model focused on syntax rather than the solution.
- ruszki 5mo agoThere is no "entire corpus of StackOverflow-esque fix suggestions" about anything which is newer than a few years. I'm using cutting edge Android frameworks all the time. Yet, LLMs fix problems even when Google/Kagi has zero answers, which happens more often than not. We are way over this requirement. I especially found that there is no difference between languages based on that. All generated code's architecture is terrible, if you don't actively manually maintain them all the time. If you don't have a few 10s of thousands of finely architected code already in your codebase, from which they can understand how it should be really done. And the reason, I think, is quite simple: the average code on the internet - regardless of market penetration of the given language - is simply bad.
- btown 5mo agoWell, for the time since then, the LLM providers have a corpus of every code suggestion they made to users and whether it resulted in positive or negative sentiment afterwards - which is arguably even more powerful. There's still some level of RLHF that's more prominent for popular languages. As you noted, of course, this doesn't apply to architecture. But that's also why I try to make sessions as turn-efficient as possible - you need every bit of context to get it to solve its own architectural rabbit holes.
- gmueckl 5mo agoTraining data can't be the whole answer. LLMs are really good at translating to different programming languages. This makes sense, given that they are derived from text translation systems. I'm getting great results in languages with comparatively small bodies of freely available code. The bigger hurdle is usually that LLMs tend to copy common idioms in the target language and if it is an "enterprise-y" language like Java or C#, the amount of useless boilerplate can skyrocket immediately, which creates a real danger that the result grows beyond the usable context window size and the quality suffers.
- lanyard-textile 5mo agoVery true. I have to steer models hard for C++. They constantly suggest std::variant :P
- socalgal2 5mo agois that bad? Godbolt got a 2x speed improvement switching from what he thought was a good fast impl to std:variant https://www.youtube.com/watch?v=gg4pLJNCV9I https://www.youtube.com/watch?v=gg4pLJNCV9I
- lanyard-textile 5mo agoMore fundamentally -- it creatively finds ways to use std::variant that I don't think any person would ever naturally conceive :) And I think it's extremely fascinating, because I have absolutely no idea why. Claude can plow through my TypeScript and read my mind easy peasy, but C++? It's just a constant battle to convince it not to use std::variant. Basically if there's any two branches of possibilities, I look forward to seeing a complicated proposal for std::variant. And then I inevitably push back, and Claude gives the "You're right, I was overcomplicating it" and it does an if/else like you'd expect. Maybe something about my particular repos cause that, but none of them use std::variant a single time. I also wonder if it's far more commonly used than I realize, that also occurs to me.
- chromacity 5mo ago
- not2b 5mo agoThat would matter if we were asking the AI to generate code open-loop: someone probably already wrote something close to what you asked for in Python. But if the agent generates code, tries to compile it, sees the detailed error messages and acts on those messages to refine the code, it's going to produce a higher quality result. rustc produces really good diagnostics. And there's a lot of Rust code online now, even if there's so much more Python and Javascript/Typescript.
- ambicapter 5mo agoLLMs don't actually semantically parse the error messages. They will generate the most likely sequence resulting from the error message based on their training data, so you're back to the training data argument.
- not2b 5mo agoThey process those error messages in the same way that they process your instructions about what code to generate. It is just more commands.
- neutronicus 5mo agoPerhaps the training data about what compiler diagnostics mean is particularly semantically rich training data.
- Tarq0n 5mo agoOf course they do, error messages get tokenized and put into the context window just like anything else. This isn't a Markov chain.
- hansvm 5mo agoExcept the presence of errors, mistakes, contradictions, and doubling-back causes LLMs to have substantially worse output, especially without dedicated sub-agents who have been instructed about that deficiency and know to process that kind of crap into better prompts to insert into a different LLM with pristine, error-free context. Without hard numbers we're both just pissing into the wind, but it's entirely plausible that the higher rate of errors matters more than the fact that those errors are more ergonomic. Anecdotally, my LLM work is a _lot_ more productive when I have it draft the thing in Python and translate it into Rust since it wastes so much time on the tiniest of syntactic mistakes.
- faangguyindia 5mo agoNo if that mattered you'd write everything in html and css. Because that has way more training data.
- weird-eye-issue 5mo agoThose are not programming languages.
- goatlover 5mo agoWASM then.
- weird-eye-issue 5mo agoThat's more of a compilation target than a programming language and I don't really see the relevancy...
- robot-wrangler 5mo ago> I could write in brainfuck with ai, but I presume, wouldn’t get the same results than if going with python. https://esolang-bench.vercel.app/ https://esolang-bench.vercel.app/
- _boffin_ 5mo agoand this sums it up right here.
- Tarq0n 5mo agoThe conclusions seem overly broad. Just because these languages are Turing complete doesn't mean they aren't massively hampered by expressiveness and amount of batteries included. To attribute all of this to training data memorization is premature.
- robot-wrangler 5mo agoOh this is a very damning paper. Using simple languages from their definitions alone is a great proxy for studying truly out-of-distribution reasoning. Also just for following simple rules/instructions correctly, because a simple enough language is practically just a grammar. This paper is terrible for anyone who wants to make the case that models can do those things well. To the extent today's AI can reason, add this to the pile of evidence that you definitely need a harness. Counter to what you hear.. that seems true for SOTA and frontier, not just toy models. Lots of people were saying many years ago someone should test exactly this, because it's obvious. Someone at megacorp probably did try and decided not to publish because they thought it was bad optics.
- gerdesj 5mo ago"I could write in brainfuck with ai" Well, go on and do the experiment! Perhaps LLMs can right code as well in BF as Python but I don't recommend it because hallucinations are really hard to notice in BF. If you are going to worry about high level computer languages and AI, you are going to have to start with getting to grips with machine code and assemblers and that. Once you know how say some Python code ends up being processed by your laptop CPU(s), then you will know when BF might be best!
- _boffin_ 5mo ago> Frontier models score ~90% on Python but only 3.8% on esoteric languages, exposing how current code generation relies on training data memorization rather than genuine programming reasoning. https://news.ycombinator.com/item?id=48100433#48102985 https://news.ycombinator.com/item?id=48100433#48102985
- th1sisoldnews 5mo ago[dead]
- bensyverson 5mo agoJust use Go. LLMs have seen a ton of it, they write it well, it compiles practically instantly, and it has all the advantages of a typed compiled language. I created a big Python codebase using AI, and the LLM constantly guesses arguments or dictionary formats wrong. Unit tests and stuff like pydantic help, but it's better to avoid that whole class of runtime errors altogether.
- hirvi74 5mo agoBut what is the selling point for Go? I get that it is allegedly hailed to be a simple language with basically no batteries included, but why is that a selling point? Does Go excel at anything no other language does?
- deleted 5mo ago[deleted]
- chickenman_98 5mo agoI think that’s sort of the selling point no? It’s really boring. It has like -10 keywords, compiles insanely fast, and has a concurrency model that’s easy to use and read. LLMs are great at using Go tooling to sanity check along the way. It’s easy to write shitty Go but it’s really pleasant to work with if you find those things compelling.
- khimaros 5mo agodon't you worry about garbage collection?
- camdenreslink 5mo agoIf you were using Python, then probably not.
- bensyverson 5mo agohaha exactly. I’m coming from Swift, and I don’t want to go back to manually releasing objects like I used to in ObjC, let alone reason about lifetimes.
- Eridrus 5mo agoThe LLMs are actually worse at generating Python than other langs, hypothesized due to quality of training data lol. I still read the generated code, so I'm not quite willing to give up on Python yet though.
- tengbretson 5mo agoAdmittedly, I have very little experience with LLM-assisted Python. However, based on the severe degradation in output quality I have seen from an LLM working with plain JavaScript as opposed to TypeScript, I can't imagine choosing to start a project in Python at the moment.
- fwip 5mo agoIt does seem like LLMs write better Python when told to use type annotations, especially when coupled with a linter.
- aix1 5mo agoI've been coding professionally in Python for about twenty years (alongside, at different times, a dozen or so other languages). I find that Claude can write great modern Python more or less out of the box, with minimal style guidance from me. I do have to nudge it from time to time to not do silly things, but overall it's really rather good.
- bluegatty 5mo agoThere's enough training data on the other langs.
- gertlabs 5mo agoSurprisingly, LLMs are actually much worse at reasoning in Python than other common programming languages for agentic coding tasks. Data here: https://gertlabs.com/rankings?mode=agentic_coding https://gertlabs.com/rankings?mode=agentic_coding
- hooloovoo_zoo 5mo agoMm, the code is constrained to run inside a game 'tick'?
- BariumBlue 5mo agoHah, I was just thinking that Python likely has a vast ocean of training data, but it's likely of lower quality, being much of it is written by beginners and those who aren't primarily programmers.
- topham 5mo agoThere's a broken idea that AI know Python because they're written in Python. Not how any of it works.
- gertlabs 5mo agoWhile recent models are capable of generalizing to any language at this point, I do think there are weights from their pretraining corpus that still leak through into how they create their responses. We observed similar language performance patterns across models from different providers, btw.
- dasyatidprime 5mo agoNot what anyone was talking about. Training corpus ≠ inference engine.
- stefanfisk 5mo agoReminds me of the time I asked Claude to write some Wordpress code for me. The results were…rough.
- 5mo ago
- onlyrealcuzzo 5mo agoI built a programming language, and LLMs can code phenomenally well in it. I don't think the training set matters that much, since there's no way they have my language in their training set! Programming languages have a lot in common. Python is kind of odd when it comes to languages.
- zuminator 5mo agoIf the training data is basically irrelevant, then an LLM should be able to iteratively improve the programming language it uses, resulting in a custom language optimally designed to maximize its own coding ability. The source code might not even be human readable natively, just translated into pseudocode on an as-needed basis.
- onlyrealcuzzo 5mo ago> If the training data is basically irrelevant, then an LLM should be able to iteratively improve the programming language it uses, resulting in a custom language optimally designed to maximize its own coding ability. I won't be surprised if one day they do. At least in their current form, I don't think they can independently design a language that is so much better than other available ones that it makes sense for them to use it. There's a very good language for almost every use case already, designing one better than the ones already available is a VERY tall order. It's almost like these languages aren't designed by morons, but built by teams of geniuses over a decade instead. It's taken me 6 months of heavily steering an LLM to build a language that is not yet even ready for production use. Maybe I'm the one slowing the LLM down. But it certainly does not seem that way. The key to a good language for them - from my experience - is maximum expression plus minimum global complexity. Anything that makes you manage memory lifetimes & memory safety is inherently unfriendly to LLMs - that's globally complex. All scripting languages allow spaghetti aliases that let you hack your way into oblivion - and LLMs gladly ride that gravy train to hell. Rust excels here, because it prevents the worst and is WAY more expressive than most people think. Go has arguably the best runtime ever built, but it's not very expressive, and it still has a lot of problems from C and scripting languages - I don't think these types of languages will be the ones people chose to write code with for LLMs in the future.
- mountainriver 5mo agoI loved from writing all my code with LLMs from Python to Rust. I’ve seen absolutely no difference, most of the time I couldn’t even tell you which it’s writing in. My programs are faster and more reliable than they’ve ever been.
- bmitc 5mo ago> Read the first few comments and surprised I didn’t see it, but training data. The voluminous amount of Python in the training data. That's actually part of the point. Almost no one writes types for Python and has complete type compliance. So all that training data is people just yoloing Python, writing a bunch of poor code in it. I honestly can't believe any experienced software engineer would decide to build systems in Python these days.
- ocschwar 5mo agoSeems to me these LLMs have a critical mass of Python training data and Rust training data, so there's no advantage for Python there. So as the article points out, an iterative process that catches the mistakes at compile time is much more suited for an AI than one that catches them at runtime.
- markboo 5mo agothat's right, we dont need to care about a lang, same as we dont care about Map when FSD promise its already end to end optimal one.
- osigurdson 5mo agoI wouldn't say I get worse results with Go than I do with Python.
- impulser_ 5mo agoPeople really need to stop assuming more training data the better. This is not how it works. LLM thrive off consistency. Go for example has significantly less training data than Python, but LLMs are the best at it. Why? Go is often written the same. You go from project to project and the code looks all the same. There only a very few ways to write Go.
- dillon 5mo agoI had an itch to give Perl another go after a 5 year hiatus. I wanted a super simple way to spawn a proxy I was building in Go, along with writing various integration tests. I used Claude Code to write the bulk of it and found Claude to be remarkable good at Perl. I told Claude to only use what’s built into Perl’s standard library rather than reaching for anything in CPAN. Turns out everything from HTTP clients, TLS and JSON are all builtin which makes it a very stable and easy way to replace what I would normally have implemented in shell scripts. My theory is because Perl hasn’t changed all that much and has a ton of training data that Claude is actually quite good at Perl for cases where you might think to write shell scripts.
- deleted 5mo ago[deleted]
- hiAndrewQuinn 5mo agoMany are saying this! https://til.andrew-quinn.me/posts/llms-make-perl-great-again/ https://til.andrew-quinn.me/posts/llms-make-perl-great-again...
- imhoguy 5mo agoPlus Perl has very efficient minimal syntax, with "Perl golf" training set it is almost like ascii bytecode for LLMs.
- obelos 5mo agoI'm not sure it's really from the lack of change, though. I've used Claude, Kimi, and other LLMs to write a ton of Perl that's jacked up with weird sugaring packages like Moose and Function::Parameters with reified types, and they seem to pick up the new idioms pretty effortlessly. It's a really unexpected fluency, frankly.
- te_chris 5mo ago1) the models do generalise so concepts translate 2) languages with more opinionated semantics and a better, more coherent community seem to be better. Python is a broad shitshow with multiple ways to achieve the same thing. Elixir is tight and focused. Claude is much better at elixir.
- imron 5mo agoLarge volumes of training data is a blessing and a curse, especially when you consider who wrote it.
- aaa_aaa 5mo agoFor some people reducing infra costs matter. Python is very very slow, even if it uses native libs.
- krzyk 5mo agoWith AI it is important to catch errors/hallucinations early, static typing helps with that. So languages with dynamic typing might hide some errors until runtime, static typing one could catch that during compilation. With dynamic ones you need way more tests to cover some of the scenarios that compiler does for others. And there is significant amount of code written "for ages" in languages that were there longer, like C, C++, Java (yes, I know that python is quite old, older than Java - 1991).