9 ms·
> Go is absolutely one of the best programming languages for LLMs for the reason you say, and Python is just what LLMs like to use to write short throwaway scri
by boxed 2mo ago
> Go is absolutely one of the best programming languages for LLMs for the reason you say, and Python is just what LLMs like to use to write short throwaway scripts.
And yet this article has pretty strong empirical data to show that your intuition here is incorrect. You should back up your statement with something more than vibes.
- win311fwg 2mo agoBest to read the comments before replying. The article is about correctness, while the parent is talking about output quality.
- boxed 2mo agoWhat is output quality without correctness? That seems like a distinction without a difference. Is the claim that LLMs produce Go code that is superficially nice looking but in fact fail to solve the stated problem? Because that's an anti-Go position I'd say.
- win311fwg 2mo agoCorrectness is binary, while quality is not. Correctness is a suitable property to act as a multiplier in your formula, where incorrect is 0 and correct is 1, but you also need other facets to find a quality gradient.
- boxed 2mo agoIn this context it's not binary. Context is everything. If it was binary there would only be 0 and 1 on one of the axis in the graph. That's not the case.
- win311fwg 2mo agoCorrectness is binary even in context. The axis of which you speak shows distance; essentially how close the programs were to being correct. Every single sample was incorrect. Think of it as being like a road trip. Arrival is binary. You have either arrived or have not arrived, but distance can tell you how close you are to arrival. Being almost there does not imply that you have arrived, however. Same applies here. Some samples were closer to being correct than others, but none were correct. They were all incorrect. But as you alluded to earlier, I don't suppose anyone wants code that fails to solve the stated problem. Given that we have empirical evidence that LLMs cannot produce correct code within a given set of problems (and I suspect that extends to most problems), correctness is a weak signal. What is a useful is to know is how much additional effort is required to make the program correct. That is what quality has traditionally meant as it pertains to code. High quality codebases are considered high quality because the effort to reach and maintain correctness is considered to be low. We do not have enough context to know for certain if that is what was meant in the earlier comment, but it seems likely. What we do know is that the comment is about quality while the article is about correctness.
- boxed 2mo agoI think "quality" here refers to "idiomatic go code", not the distance from the end broken state to a finished product. At least that's how I read it.
- win311fwg 2mo agoYou seem to be mixing up the different things we are talking about. Yes, quality is what the earlier commenter is talking about. Yes, idiomatic code is often considered to be what makes code high quality because idiomatic code is believed to make it easier to reach correctness (easier to read, easier to reason about, easier to test, etc.) than if you have to work with "spaghetti", which is considered low quality because it makes it much harder to ensure the code is correct. Quality exists on a gradient. There exists a full spectrum between a complete spaghetti monster mess and perfect architected idiomatic code. This we both read the same way it seems. Whereas the article is about correctness. The program is either correct or not. Distance was used in the article to share how far away a given program was from being correct. But distance and quality not the same. This is a completely different way to look at software as compared to the topic of quality, as we already established at the beginning of our discussion.
- boxed 2mo agoI think that idiomatic Go code is not in fact easier to read or reason about on the business domain level, because it's too low level and fiddly. That's why we don't write in assembly anymore. Assembly is much more "easy to reason about" per line of code than Go, but it's also terrible on the business domain level for the exact same measurement.
- win311fwg 2mo agoSo you don't believe in "spaghetti" code? No matter how a codebase is written you will consider them all to be equally readable, equally able to reasoned about, equally able to tested, etc.? You see no difference between "organized" code and one giant 200,000 line chaotic function? We mostly avoid writing assembly because it isn't portable. So-called "portable assembler" is still very popular. One of the most used languages out there. That has little to do with the quality of codebase, though. Focus, my man.
- YuechenLi 2mo agoIf you actually read the data, especially the distribution graph in the last image, the conclusion that it draws is "the run-to-run distribution variance is so big that there doesn't seem to be a correlation that can be drawn from this experiment", pretty much every language has similar-ish distribution ranging from ~20 to 34, and Clojure is only the worst because GPT has a tendency to write code that contains a particular byte manipulation mistake that it repeatedly makes, not that GPT is bad at Clojure or anything. My experiences are of course anecdotal, but if you have some other strong empirical data to show, I'd love to see it.
- boxed 2mo agoYea exactly. The article says there's too much noise to make any conclusion. You made a conclusion that there was a strong signal. Those things are opposites.