5 ms·
Hah, I was just thinking that Python likely has a vast ocean of training data, but it's likely of lower quality, being much of it is written by beginners and th
by BariumBlue 5mo ago
Hah, I was just thinking that Python likely has a vast ocean of training data, but it's likely of lower quality, being much of it is written by beginners and those who aren't primarily programmers.
- topham 5mo agoThere's a broken idea that AI know Python because they're written in Python. Not how any of it works.
- gertlabs 5mo agoWhile recent models are capable of generalizing to any language at this point, I do think there are weights from their pretraining corpus that still leak through into how they create their responses. We observed similar language performance patterns across models from different providers, btw.
- dasyatidprime 5mo agoNot what anyone was talking about. Training corpus ≠ inference engine.
- stefanfisk 5mo agoReminds me of the time I asked Claude to write some Wordpress code for me. The results were…rough.
- dariusj18 5mo agoThat was the hardest part of learning PHP, all the code examples online were just awful.
- andai 5mo agoWorked on a PHP project once. Every time I asked why something was done a certain way the answer was "dunno, we copy pasted this code snippet." Certain popular PHP codebases appear to use a similar methodology.
- Sohcahtoa82 5mo agoIt's why I consider PHP to be "RCE as a Service". So much copy/pasted code, some of it REALLY bad, and PHP has a lot of foot-guns that can lead to RCE.
- FireBeyond 5mo agoAll my vibe coded projects (personal) are Go backend services, with Typescript/React frontend. And my thoughts were based on similar things. Like why I wouldn't use PHP for that, either.
- librasteve 5mo agoI was (pleasantly) surprised by Claude Code doing Raku - also with a limited training set (~2000 Stack Overflow, a bunch of Rosetta, 2,500 modules). I put this down to the quality of the code for the core community who are all frankly uber-gremlins.
- polytely 5mo agoYeah Raku feels so expressive and lovely to me with the help of an AI assistant. I've only done toy programs and scripts with it but it is actually so nice.
- kraf 5mo agoThat's what I'm thinking too. There is a lot of noise and I know teams where the majority of the people writing Python just have no idea what they're doing. I'm working with Clojure which is used mostly by senior engineers and it still blows my mind how well Claude writes software in it even though it's a fringe language. It's even able to pick up in-house DSLs written with macros.
- chamomeal 5mo agoWas looking for somebody to mention that LLMs are weirdly good at clojure! Have you tried hooking your LLM up to a repl? It’s crazy stuff. Type systems are a great feedback loop, but a LLM with a REPL is something else. Much more dangerous though!!
- smoe 5mo agoHaving used Python on and off for 20 years, my experience with LLMs writing Python has been mixed. I don’t think that’s necessarily because of a low-quality dataset, but rather because Python’s applications are so broad and the language has gone through several paradigm shifts over time: sync vs. async, typed vs. untyped, scientific Python looking very different from web application code, some people really wishing it were an FP language, and others doing the clean-architecture OOP onion soup. It has gotten so fragmented. Recently, I had a more pleasant experience using LLMs with Go. It reminds me a bit of Python 2.x, when the community seemed, in my view, more focused on embracing a stupid simple language, with everyone trying to write roughly similar "Pythonic" code.
- stingraycharles 5mo ago> Having used Python on and off for 20 years, my experience with LLMs writing Python has been mixed. I don’t think that’s necessarily because of a low-quality dataset, but rather because Python’s applications are so broad and the language has gone through several paradigm shifts over time If there’s one language that is the prime example of this, it’s C++, and according to this benchmark it ranks incredibly high. I’m also thoroughly confused why Kimi 2.6 scores 83% while Opus 4.7 scores 67% for C++, GPT5.5 isn’t even in the top10. Gemma 4 31B scores 100% success rate for Python (!!) while Opus 4.6 only 65%. This benchmark really seems to be all over the place and doesn’t make sense.
- gertlabs 5mo agoThe more filters you apply (single model and single language, especially if you also filter by pipeline like agentic vs one-shot), the fewer samples, so there is variance. Known limitation that is inevitable with any finite budget. This is why we are selective about adding more languages because it will dilute the amount of samples we can run per language per model. But the aggregated statistics hold up well and are very consistent in our testing.
- stingraycharles 5mo agoI just applied a single filter, programming language.