4 ms·
I don't have a dog in this fight, but a few points that look a little suspicious: - The release with the highest number of attributed bugs is the release _righ
by aesthesia 4mo ago
I don't have a dog in this fight, but a few points that look a little suspicious:
- The release with the highest number of attributed bugs is the release _right before_ the first release with Claude-coauthored commits, released in January; is there a chance that unattributed LLM-authored commits made it into this release?
- The release attribution methodology is not great, since it will tend to attribute bugs introduced in a minor version update to the longest-lived patch release of that minor version. I doubt that 3.4.1 actually introduced a lot of bugs, but since it was released a day after 3.4.0, bugs that were introduced in that release get attributed to 3.4.1.
- Relatedly, more recent releases have had less time to have bugs filed against them, so there may be a bit of a bias toward evaluating recent releases as less buggy.
- logicprog 4mo agoYour first and second points seem to contradict each other because if all of the bugs for 3.4.1 should be attributed to 3.4.0, that pushes the timetable back even further that unattributed LLM commits would have to have been being committed to the project, which just makes your point even more absurd. Which brings me to my overall response, which is that there is absolutely no evidence, and nothing even intimating this hypothesis, that LLM commits were secretly being added to earlier releases before they were attributed, and that's why the rate of bugs is higher. There's no reason to think that it's an unreasonable thing to think, and there's no evidence for that whatsoever unless you beg the question and assume that higher bug counts must automatically indicate AI involvement, which is just circular reasoning. You're essentially just making up a hypothesis out of thin air to preserve your point. Regarding your third point, that one's fair, but I've done the analysis and I can put it up if you want, as to how long it usually takes to find bugs and how far through the release cycle we are for each version.
- jonquark 4mo agoIsn't the metric that you've used "bugs per commit ~ per new line of code" going to miss the issue? All code is technical debt. If rsync releases used to have 500 lines changed and 5 bugs in and AI-powered rsync releases have 50000 lines and 500 bugs, it's the same bugs/line but much worse experience for the user? I've not looked into the details of this case and I do use AI assistance coding at work but in my experience, the problem is that it's too easy to write lots of code and therefore hard to review the huge volumes of code and this analysis will ignore that? edit: actually your table shows there weren't unusually large numbers of commits in this release, so perhaps my initial skepticism shows a bias I have?
- tolciho 4mo agoOpenBSD used to have sqlite in base, but the code churn rate was too high to review. This was well before the recent LLM craze, so a human (perhaps not a normal one, though) already sufficies to generate too many changes for others to check for errors.
- aesthesia 4mo agoSorry, I should have said this explicitly in the original comment: I think you're likely _correct_ that there isn't a clear increase in the rate of bugs attributable to LLM-authored code in rsync. Your analysis provides evidence in this direction; these are just the things that made me go "hmm". They're not accusations or claims that the conclusion is invalid. But they're definitely things to be curious about. Regarding unlabeled LLM-authored commits, I don't think it's unreasonable in general to think that an open-source project might have had unlabeled LLM-authored commits at some point before 2026. Looking more closely at rsync's recent commit history, I think it's less likely in this case. There's just a low number of commits in general, _until_ large batches of Claude-authored commits start showing up early this year. But this then raises some questions about the bugs-per-commit metric; it does correct for something like "size of release", but also obscures a significant shift in commit velocity that may be downstream of adding LLM development tools to the workflow. Like I said, I don't have a dog in this fight, and I try not to approach sorts of questions from a position of explicit advocacy. I do think it's an interesting question, though, and we should try to understand what the data is actually telling us.
- OptionOfT 4mo agoYou can use LLMs in multiple ways, from very hands on to make local changes to completely hands-off. I've seen plenty of code that was LLM generated but the commit message itself did not have the co-author attached to it. This only seems to happen when someone's interface to the codebase is completely though Claude/Codex/..., and those are usually the most verbose commits, and yet they say the least, because they just summarize the code changes, not the why. On the other hand I've seen developers using Claude as a tool. They have VSCode open and a terminal window with Claude and go back and forth, ensuring they write correct code, and leave the plumbing to Claude. So maybe the author of the code started off small and it grew over time?
- hparadiz 4mo agoI would expect a mature code base like rsync to have a lot of unit tests and integration tests and frankly if there's not enough that such bugs haven't been caught; that should be your first use of LLMs in order to setup some deterministic guidelines when you do start making changes to your actual code. I have been experimenting with both aforementioned styles with interesting results.
- cyanydeez 4mo agoI've had a local LLM spending weeks trying to write tests. then debug those tests. then write antipatterns and patterns for those tests. It's amusing. It's not terrible, but tests arn't going to save you from a malicious tester.
- duskwuff 4mo ago> I would expect a mature code base like rsync to have a lot of unit tests and integration tests You might be surprised. C applications which interact heavily with the system - like rsync - can be tricky to test comprehensively, as it's nontrivial to inject faults into system calls. If the application is architected to support this kind of testing, or uses a HAL, that may make matters easier - but an older codebase like rsync probably isn't.
- PunchyHamster 4mo agoLet's start with most outright alarming error - the claude statistics are taken out of whole 2 data points
- logicprog 4mo agoThat's sort of the point. There isn't enough data to extrapolate, and yet that's exactly what those outraged about AI were doing, and when you do do the very minimal types of analyses (permutation tests, and looking at distributions, mostly) that are actually valid, safe, standard, and useful to do on such low amounts of date, again, no evidence for the outrage shows up, and the two releases look so normal that it sort of shows no one would've cared if they hadn't known or found out that Claude was involved. I really think this a much better standard of evidence — limited though it is — to outrage-fueled cherry-picked anecdotes, which is what has been driving this whole thing. If you disagree, and think the outrage should go one when I've shown there's an absence of evidence entirely for it (although of course, that's not evidence of absence; maybe I'll have to eat my words 5 releases down the line, but appealing to that now feels like a Russell's Teapot), would you care to explain why?
- ofjcihen 4mo agoI know you’re defending your work here but this behavior does absolutely nothing to help your point.
- logicprog 4mo agoFair point. Let me edit (if I still can) to tone it down.
- PunchyHamster 4mo agoyou could've literally just waited few more releases, but no, have to catch the hype wave before news are cold > that are actually valid, safe, standard, and useful to do on such low amounts of date, if you presented paper with that amount of data points you'd be laughed out of the room
- 4mo ago
- theteapot 4mo agoAgree. From the article: > Here's my favorite part, though. Digging into the data, one of the first things that jumped out at me with blinding clarity was that the worst release, by far, in rsync history was entirely prior to the introduction of Claude ... And yet nobody noticed. Language really does suggest the article's author does have a dog in this fight and is cloaking opinion in fancy statistics jargon. "Blinding clarity"? All you have to do is draw a plot. And anyway, v3.4.1 was 2025-01-16, technically well within the AI assisted coding era and before attribution was becoming standard practice.
- iandinwoodie 4mo agoAlso from the article: > "Claude clearly made things worse" &emdash; the main claim This article was clearly generated by AI, yet I found no mention/attribution of that by author. How likely is it than someone who vibe codes articles would also vibe code the underlying analysis and be eager to accept an outcome that is highly validating of that person’s workflow? I’d say very.
- int_19h 4mo agoAre the numbers wrong? That's the only relevant thing here. Also, humans do use em dashes, just FYI.
- davrosthedalek 4mo agoYes, I do for example. And the author discussed the use of AI pretty exhaustively in point 0 of the post.
- deleted 4mo ago[deleted]
- latexr 4mo ago> Are the numbers wrong? That's the only relevant thing here. Data without interpretation is irrelevant, and correct numbers can be interpreted wrongly, either on purpose or by mistake. I’m not saying any of that happened here, only that “are the numbers wrong” is not the only thing that is relevant. > humans do use em dashes, just FYI. Your parent comment is not complaining about em dashes, they are pointing out the article has a literal “&emdash;” in it.
- hariseldom 4mo agoI started to look into the same thing considering releases are quite infrequent. To avoid the issue of unattributed LLM-authored commits, in my opinion the analysis should include a comparison to bug severity before and after release v3.3.0 (date April 6th, 2024)