15 ms·
Did Claude increase bugs in rsync?
- wookmaster 4mo agoClaude is just a tool ? The developers who merged that code and didn't properly test increased the bugs.
- everdrive 4mo ago"Did cars increase traveling deaths?" "Cars are just a tool. The drivers who piloted the vehicles and weren't careful enough [are responsible for the deaths.]"
- Angostura 4mo agoThis tool is claimed to be able to find and fix bugs.
- roywiggins 4mo agoIf something's a bad tool that misleads people into doing bad work, it would be good to know that.
- ebiederm 4mo agoPlease read the article. The unsolicited security reports are the issue.
- runarberg 4mo agoFeels like something a bad (and potentially dangerous) tool would say.
- rovr138 4mo agoI'm just curious about testing. Is this a configuration that's not common and thus not tested? If people think they can do better, I want to see their forks and them keeping up with it. https://github.com/RsyncProject/rsync/graphs/contributors?from=6%2F1%2F2024 https://github.com/RsyncProject/rsync/graphs/contributors?fr...
- the_real_cher 4mo agoIs there a non vibe coded fork of rsync?
- throwaway7356 4mo agoYes: https://news.ycombinator.com/item?id=48390931 https://news.ycombinator.com/item?id=48390931 So far it reintroduced several security issues and replaced the README.md.
- MYEUHD 4mo agoThere is openrsync, which the OpenBSD re-implementation of rsync It's not a fork, but it's 8 years old, and is already shipped by default in OpenBSD and macOS.
- logicprog 4mo agoTo quote Tridge: > As to all the people saying “I’m going to package openrsync for platform XXX and we’ll use that!”. I find that rather amusing. If you do decide to go down that path I’d suggest you try the new rsync test suite on openrsync if you can stomach something that an AI has helped write. I tried it today and openrsync currently fails 85 of 98 tests, so I’m sure it won’t take you long to get it up to speed. You run it like this “./runtests.py — rsync-bin=../openrsync/openrsync — use-tcp”. Admittedly a lot of the failures are just features openrsync doesn’t have, but still, it’s not a great result.
- MYEUHD 4mo agoI have already been using openrsync even before the recent AI drama. Just like I have been using doas for several years. All I need is `rsync -urvP` and I suspect the majority of users don't need the advanced features either. The smaller code base also means less bugs and vulnerabilities. As an example doas is ~1k lines vs 160k for sudo. That surely means a smaller attack surface. The same is true for openrsync and rsync at approximately 18k vs 57k lines.
- 4mo ago
- nairboon 4mo agoIs this an analysis made by/with Claude?
- quentindanjou 4mo agoIt very obviously is. "The Outlier Nobody Noticed" -_-"
- overgard 4mo agoFWIW, I asked ChatGPT to review the article just for my amusement. It's conclusion was: "My honest assessment is that this is a competent calculation performed on a badly confounded measurement, followed by conclusions substantially stronger than the calculation warrants. It is useful as a rebuttal to “the Claude releases are obviously unprecedented disasters,” but not as evidence that Claude was harmless."
- Polarity 4mo agoso the answer is: no. actaully less bugs. thanks
- gjvc 4mo ago"fewer"
- deleted 4mo ago[deleted]
- davrosthedalek 4mo agoFirst rsync and now less? What comes next, cat?
- gjvc 4mo ago/bin/true
- davrosthedalek 4mo agoRight, that's now /usr/bin/true !
- deleted 4mo ago[deleted]
- geraneum 4mo ago> But the critics' accusation is also blunt: "Claude is making things worse." A blunt instrument is the fairest response. So the criticism was bad, and that somehow makes it ok to use a bad metric?
- logicprog 4mo agoThat's not what I'm saying. What I'm saying is that if the criticism is referring to a broad set of metrics like bugs per release and number of commits that were made by Claude, then it's correct to look at precisely those things because that's what the claim is about.
- deleted 4mo ago[deleted]
- abirch 4mo agoAI + Interest != Expertise I come to hn because I get very nuanced, informed information and glorious puns.
- epolanski 4mo agoWhat would be a better one?
- faitswulff 4mo ago> The analysis uses a single metric: bugs per 10 commits (bugs/10c). Bugs per commit as a metric papers over severity, both in terms of security severity as well as the effect on the user. A mislabeled button has the same weight as the entire app crashing in this framework.
- skeledrew 4mo agoThere was no analysis of severity in all of the rage posting that occurred. The single point being pushed was "use of an LLM led/leads to more bugs". The author specifically states that's what they're addressing (blunt accusation -> blunt response).
- atmavatar 4mo agoThe specific problems mentioned were all reasonably severe. The original post itself described a show-stopping bug: So my systems recently updated to rsync 3.4.3, and as soon as that happened my backup system - which does incremental backups using multiple --compare-dest= arguments - started to fail on anything but a full backup. Incremental backups is perhaps the primary use of rsync, and they were broken for this person. That's pretty severe. The second reply is similar: i wondered why my 3d printers were running like sh*t and at 100% cpu; turns out log2ram uses rsync. This one I took with a grain of salt, since it read more like a dogpile than an actual bug report. However, if it's genuine, it's also reasonably severe. Later in the comments, someone attempted to provide a list of issues that had been added: https://github.com/RsyncProject/rsync/issues/929#issuecomment-4585695638 https://github.com/RsyncProject/rsync/issues/929#issuecommen.... The list included several failures to build or run rsync that appear to have resulted from broken backward compatibility. That seems reasonably severe. If intentional, I would have expected mention in the release notes about the removal of backwards compatibility, but none was made. The issue comments already degraded into a lot of unnecessary vitriol even before the above mentioned comment and only gets worse from there, so I stopped. But, the fact remains that the whole issue started with a severe bug. I applaud the attempt at dispassionately analyzing whether the recent LLM releases of rsync were normal or outliers as far as bugs are concerned, but I don't think you can do so properly without analyzing severity.
- gadrev 4mo agoOk. $ apt-cache policy rsync | grep Installed Installed: 3.4.1+ds1-7ubuntu0.2 $ sudo apt-mark hold rsync rsync set on hold.
- logicprog 4mo agoDid you face any actual bugs or regressions? Or are you doing this just because of the bandwagon that's going around right now? Because until you can actually present an argument for why this release is worse than any of the others, which is precisely the subject of my post, then this is not an argument against my post at all. This is just a self-referential appeal to authority.
- gadrev 4mo agoNah, I skimmed TFA but then I went into the linked GH issues thread, and that's the one that scared me a bit. I just want to hold it for a while and not run into some of the things I'm reading since I'm on the latest ubuntu. Just a precaution. I didn't have the time to actually think about any "arguments" at all tbh it's just a knee jerk reaction as I get ready to log off for the weekend. Not actually looking to argument for or against your post at all lol.
- overfeed 4mo ago> Did you face any actual bugs or regressions? This is a terrible argument; I didn't need to have had secrets exfiltrated before applying row-hammer mitigations. If rsync is the cornerstone of my backup strategy, and has been for years, I need to trust that on its correctness, and for it to not lose my data. If I wait until I "face any actual bugs or regressions" - that will be far too late. Stability is another issue not discussed. If the error rate holds steady, but number of significant PRs merged per release goes up from 5 to 200, that would be huge net-negative for my use case.
- imurray 4mo agoThat version has security fixes from the same day as the latest rsync release: https://ubuntu.com/security/notices/USN-8283-1 https://ubuntu.com/security/notices/USN-8283-1 As usual, Ubuntu backported fixes and didn't upgrade to a new version. Whether or not they also backported regressions in edge cases that afflict the latest rsync, I don't know. Pinning the Ubuntu package may prevent getting further regressions, but is preventing you getting any future such backported security fixes.
- scsh 4mo ago> It does not control for commit complexity, security intensity, or bug severity. It does not distinguish between a one-line typo fix and a CVE patch. It is a blunt instrument. But the critics' accusation is also blunt: "Claude is making things worse." A blunt instrument is the fairest response. If by fairest you mean to say that this analysis and response is sufficient, then I'm sorry but I have to disagree. We really need to understand if the nature of the bugs are worse from a user's perspective. Even if the rate stayed unchanged, if the result is the perceived quality of the software declined then I would personally consider that worse, especially if I were a project maintainer. That's not meant to be wholly dismissive either. But in general, I don't think quantitative analysis alone is enough to fully answer this type of question.
- skeledrew 4mo agoBut it is fair. Up to this point I have yet to see anyone say they did an analysis of the code and found X regressions of Y severity. All they say is "there are more bugs because LLM". This analysis, which you can verify yourself if you wish, says "the bugs [number of] are pretty average even with LLM", which is a direct response to that. If you'd like a more nuanced analysis you're welcome to do one and share the result, if you're so inclined.
- MostlyStable 4mo agoThat which is asserted without evidence can be dismissed without evidence. This is more evidence, and of greater rigor, than was used to make the assertions. That's good enough for me. If someone wants to actually do the work to support the original claims with better evidence, great. I'd love to see it. Until then, I'm going to not worry about this issue.
- ex-aws-dude 4mo agoThe burden of proof is on the one making the claim?
- logicprog 4mo agoOkay, I really have to point out to everyone: the numbers and report cards are TEMPLATED IN BY A SCRIPT. Hallucinations are a moot point. https://github.com/alexispurslane/rsync-analysis/blob/main/scripts/regression_report.html https://github.com/alexispurslane/rsync-analysis/blob/main/s...
- pushcx 4mo agoWhat followed was extraordinary: 329 comments and counting, ranging from thoughtful concern to outright harassment. The thread did not stop at words. One user posted My Little Pony drawings of themselves strangling the "project janitor that pushed vibecoded commits": It spread to Hacker News and Lobsters, generating hundreds more comments. This is false, it did not appear on Lobsters. Here is the function in the codebase that prohibits this kind of brigading: https://github.com/lobsters/lobsters/blob/main/app/models/story.rb#L371-L407 https://github.com/lobsters/lobsters/blob/main/app/models/st... Please correct your article.
- logicprog 4mo agoI have done so! that was a misremembering on my part. first mention of Lobsters is now here: > On Lobste.rs, in response to the Medium essay Tridge himself posted in response, finally some users like boramalper begin to actually ask for evidence one way or another:
- pushcx 4mo agoThanks, I appreciate you sorting out the timeline on such a heated issue.
- tptacek 4mo agoIt is neat that Lobsters has this feature (and HN should too), and I'm glad you took a beat to explain it. I think you didn't need the last sentence, though.
- thorum 4mo agoUnfortunately for the people mad about this, I predict the only thing they will accomplish by pressuring the rsync maintainers, is to discourage everyone else from responsibly disclosing their use of AI. You’re just going to make people disable Claude attribution on their commits to avoid drama.
- potsandpans 4mo ago[flagged]
- automatic6131 4mo ago"let's go the opposite way" Do you have any popular open source projects? Or are you just an Internet gremlin?
- potsandpans 4mo ago[flagged]
- deleted 4mo ago[deleted]
- elnatro 4mo ago[flagged]
- matheusmoreira 4mo agoIt makes no sense at all to do that. The only thing that matters is whether the code is good.
- potsandpans 4mo ago[flagged]
- eschaton 4mo ago
- dang 4mo ago[stub for offtopicness] [see https://news.ycombinator.com/item?id=48416020 https://news.ycombinator.com/item?id=48416020 for how all this happened in the first place]
- perching_aix 4mo ago[dead]
- duk3luk3 4mo agoThis article is unfortunately unreadable because all of the prose is unfiltered LLM slop.
- roywiggins 4mo ago> A simple distributional analysis of every rsync release with bug data. No model. No assumptions. Just placement. If you want me to read your analysis, you are going to have to make it not read like Claude wrote it. What does "placement" even mean here?
- logicprog 4mo ago"Placement" as in where the Claude-driven releases exist within the existing distribution of bugs per 100 commits. If they're not OOD, then nothing is unusual. Also, it wasn't written by Claude FWIW, GLM 5.1.
- rroblak 4mo agoYeah, made me chuckle that an LLM— probably Claude— was used to write this. The use of "regime shift" is what gave it away for me. I've never seen a human write that, but Claude does from time to time. At least they removed occurrences of "load-bearing".
- roywiggins 4mo ago"quietly" seems to be the new one recently
- overgard 4mo agoThe TLDR seems to be: needs more data.
- aesthesia 4mo agoI don't have a dog in this fight, but a few points that look a little suspicious: - The release with the highest number of attributed bugs is the release _right before_ the first release with Claude-coauthored commits, released in January; is there a chance that unattributed LLM-authored commits made it into this release? - The release attribution methodology is not great, since it will tend to attribute bugs introduced in a minor version update to the longest-lived patch release of that minor version. I doubt that 3.4.1 actually introduced a lot of bugs, but since it was released a day after 3.4.0, bugs that were introduced in that release get attributed to 3.4.1. - Relatedly, more recent releases have had less time to have bugs filed against them, so there may be a bit of a bias toward evaluating recent releases as less buggy.
- logicprog 4mo agoYour first and second points seem to contradict each other because if all of the bugs for 3.4.1 should be attributed to 3.4.0, that pushes the timetable back even further that unattributed LLM commits would have to have been being committed to the project, which just makes your point even more absurd. Which brings me to my overall response, which is that there is absolutely no evidence, and nothing even intimating this hypothesis, that LLM commits were secretly being added to earlier releases before they were attributed, and that's why the rate of bugs is higher. There's no reason to think that it's an unreasonable thing to think, and there's no evidence for that whatsoever unless you beg the question and assume that higher bug counts must automatically indicate AI involvement, which is just circular reasoning. You're essentially just making up a hypothesis out of thin air to preserve your point. Regarding your third point, that one's fair, but I've done the analysis and I can put it up if you want, as to how long it usually takes to find bugs and how far through the release cycle we are for each version.
- jonquark 4mo agoIsn't the metric that you've used "bugs per commit ~ per new line of code" going to miss the issue? All code is technical debt. If rsync releases used to have 500 lines changed and 5 bugs in and AI-powered rsync releases have 50000 lines and 500 bugs, it's the same bugs/line but much worse experience for the user? I've not looked into the details of this case and I do use AI assistance coding at work but in my experience, the problem is that it's too easy to write lots of code and therefore hard to review the huge volumes of code and this analysis will ignore that? edit: actually your table shows there weren't unusually large numbers of commits in this release, so perhaps my initial skepticism shows a bias I have?
- MagicMoonlight 4mo ago[flagged]
- Etheryte 4mo agoEmdashes don't really tell you much anything these days tbh. Many languages use them regularly and those people often bring the habit with them when they write in English — me included. Plus I would imagine every major model has tuned them way down at this point due to the backlash.
- logicprog 4mo agoI rewrote all the AI prose several hours ago with purely my own. I like em-dashes, and specifically use them with spaces as a habit. I don't know what to tell you.
- mikaeluman 4mo agoNot going to critique this survey. Must have taken a lot of time and required a lot of patience. Great work! I think it will be up to some group in academia to make a real full blown study across several repositories. There must be tons to learn on how LLMs have changed software development and perhaps the cleanest separation will simply be going by what repositories declare e.g. "No LLM involved" vs those that proudly do the opposite or are neutral. Bugs is not the only variable of interest here. I am guessing someone is already doing this as we discuss it here...
- logicprog 4mo agoAnother update: did an automated severity analysis on each bug report (~2000 of them!) using an LLM at temp=0 with a very strict rubric (and I checked to make sure that it rated things in a consistent, stable way using it). The rubric, LLM used, and some example ratings are included in the methodology section. For now, the information was just stored per-bug in the DuckDB and used to filter out non-bug bugs, to get a clearer signal. I'm going to try to use it to see if the post-Claude bugs were more severe in any way next.
- tptacek 4mo agoThis is a neat post and I'm glad it got written and this is a little bit off-topic but: Hey, 'logicprog, your writing is fine! Use LLMs to critique your writing, check its structure, vet your choice of topic sentences, check flow from graf to graf and section to section, look for passive voice and overused words. LLMs are fantastic for that. But don't use a single word an LLM suggests in your actual writing. If it suggests something really fucking good, too bad, those words are disqualified. It's an easy red line to adhere to, easier than it sounds, and it'll keep your writing human. (You ended up somewhere around here anyways, but that was after you posted something with LLM-written language because you weren't confident enough in your own writing. The things you do "worse" than an LLM are what make you you; be protective of them!)
- logicprog 4mo agoThank you!
- KronisLV 4mo agoPretty cool site! > v3.4.3 has been out long enough that its rate (5.00) is already comparable to historical releases. The "wait and see" argument is an appeal to an unknowable future that shifts the burden of proof away from the critics. If more bugs surface, they will enter the distribution like every other release. There is no reason to expect a regime break. I mean, as someone who uses LLMs, it might be a good idea to consider how one might limit the amount of bugs that will appear in the future at least a little bit: parallel iterative code review loops would probably be the easiest and most applicable to LLMs, though I guess test coverage and other code analysis tools help too.
- yobid20 4mo agoneeds a tldr; im not reading all that. maybe claude can summarize it for me.
- logicprog 4mo agoAnd anti-AI people accuse people who use AI of being intellectually lazy. First of all, it's long because it's expanded to respond to all the criticisms. It seems that either something can be short, and dismissed as incomplete, or it can be complete, and dismissed as being long. Nice Kafka trap. Additionally, there's literally an Executive Summary section right there, for your TLDR.
- noAnswer 4mo agoAsked your Clanker what a joke is.
- PunchyHamster 4mo agoThe fact last few commits were attributed to claude doesn't mean previous ones didn't use it. Also if you write a paper where you get statistical conclusions out of whole 2 datapoints you'd be laughed out of the room
- logicprog 4mo ago> Also if you write a paper where you get statistical conclusions out of whole 2 datapoints you'd be laughed out of the room I'm using methods appropriate to that low amount of data, first of all. Second of all, since I'm only trying to show there's no evidence for the anti-AI hypothesis (not disprove it, or prove the null hypothesis), that's sufficient in itself. Also, I wonder why nobody said things like you're saying ("there's too little data to tell") in response to all the absolutist claims that AI caused rsync to get worse? > The fact last few commits were attributed to claude doesn't mean previous ones didn't use it. At this point, you're just positing Russel's Teapot: you'll keep assuming more and more of the code was "secretly" Claude when there's no evidence for it and no reason to think so, just because you've started with the assumption that Claude makes things worse and you want to find a way to prove it.
- vintagedave 4mo agoWhy not? Claude marks its commit messages. That there were none, and then there were, seems a signal. Especially since if the earlier commits were so clearly AI authored yet without the Claude marker, surely you or anyone would be able to spot them. You could say, X commit does not have the Claude commit marker yet was AI written. But for all the speculation on this thread, I haven’t seen anyone actually doing that. What may be possible is that the rsync maintainers used AI to assist yet reviewed and edited themselves, as many devs do, and if so then the stats in this article are still notable: there are no poor quality outliers that can reliably be attributed to AI and if one specific release (3.4.0) was, the subsequent releases which presumably also had as much AI as this speculative hidden AI release only show improvement and thus act as a pro-AI argument. The blog has many more datapoints than two. It compares many releases. You’re looking at 2-vs, not 2.
- mwkaufma 4mo agoSmokescreen of highly-contingent analysis and appeals to authority over a premotivated-conclusion.
- logicprog 4mo ago[flagged]
- bakugo 4mo agoYour analysis was so thorough, rigorous, and objective, that you couldn't be bothered to write it yourself. Do you genuinely believe an article written by AI defending itself is going to convince anyone who wasn't already on your side? All you're doing is giving more fuel to the "anti-AI crowd" you hate so much.
- logicprog 4mo agoOkay, so you didn't respond to any of my rebuttals — like the double standard between anti-AI and pro-AI claims, one of which gets to make claims based on cherry-picked anecdotes, and the other which must produce rigorous studies — you're just going to insult me/my work. Cool. > Your analysis was so thorough, rigorous, and objective, that you couldn't be bothered to write it yourself. Do you genuinely believe an article written by AI defending itself is going to convince anyone who wasn't already on your side? Except that I did. I spend days comparing and manually deciding on metrics and methodology – I did not use the AI to decide what I would do or how I would do it, so it is not "the AI defending itself" — then refining things, adding more angles to analyze, and, as I literally say in the opening section, I rewrote all the prose in the entire document just to satisfy critics like you. That sounds like "could be bothered" to me. But people like you will never be satisfied. Also, even if I hadn't done all that work, that wouldn't make it not rigorous (it clearly is) or objective (it is as objective as it can be with so little data). You're bikeshedding to avoid the point.
- bakugo 4mo ago> like the double standard between anti-AI and pro-AI claims, one of which gets to make claims based on cherry-picked anecdotes, and the other which must produce rigorous studies This statement is honestly so ridiculous that I felt it didn't warrant a direct response, but here's one anyway: AI enthusiasts have been proudly proclaiming for literal years that AI makes them 10x as productive based on cherry-picked anecdotes with zero empirical evidence to back it up. It's way, way too late to claim hypocrisy here. As I stated under the original submission about this topic, irrational anti-AI behavior is usually just an equal and opposite reaction to irrational pro-AI behavior. > I rewrote all the prose in the entire document just to satisfy critics like you. And that doesn't help. If anything, editing the AI output to make it read less like blatant slop just comes off as deceptive, like you're trying to hide the fact that the analysis was AI generated. Looking at the commits, you were adding more AI generated text less than 2 hours ago[0] before quickly editing out one of the most blatantly sloppy sentences I've ever read[1]. Regardless, the final contents of the article are not the main issue. Even if we ignore the bias clearly on display there, the premise alone is enough to dismiss the entire thing as heavily biased and chasing a pre-determined conclusion - of course someone who is so dependent and trustful of AI that they decide such an analysis on the bugginess of AI code should itself be written by AI is going to steer the conclusion towards "actually AI code is good and you luddites are overreacting". The entire concept is so tone-deaf that failing to notice it or predict the criticism before publishing is enough to prove the bias. [0] https://github.com/alexispurslane/rsync-analysis/commit/e0293b4dd1e3af6a5d761705479631c179a0d89e https://github.com/alexispurslane/rsync-analysis/commit/e029... [1] https://github.com/alexispurslane/rsync-analysis/commit/740b3f50809b164ef7025b2ee35573fe98a2e0d8 https://github.com/alexispurslane/rsync-analysis/commit/740b...
- jrflowers 4mo agoTl;dr: Yes, it did. Here is some math showing that you shouldn’t care about that.
- logicprog 4mo agoIn what way did it create more bugs? It literally doesn't show up in the data. What are you talking about?
- jrflowers 4mo agoThe only reason why people are talking about this is because of the bugs in the code that the chat bot generated OP. People updated to a version of rsync that didn’t work right, the one with all the bot commits in it. This blog post is about how claude didn’t create more bugs than usual if you think about it in one very specific way, not that it didn’t create more bugs at all. It is like if your neighbor opens your door and a dog walks in, there’s no point in doing some weird analysis about all the times you yourself have let a dog walk in. He still did that.
- logicprog 4mo agoThat doesn't make any sense, what?
- jrflowers 4mo agoI am not sure what to tell you here, because somebody literally posted the code of a bug that Claude inserted here https://news.ycombinator.com/item?id=48419197 https://news.ycombinator.com/item?id=48419197 And your response to someone pointing out that sloppy, buggy code that Claude introduced, was to just quote Tridge (which does not in any way refute the fact that you’re looking at a bug that Claude introduced to the code) https://news.ycombinator.com/item?id=48419621 https://news.ycombinator.com/item?id=48419621 I’m not entirely sure what the purpose of this project is (maybe to “prove” Tridge’s opinions about LLMs and human intelligence that he made in the linked blog post to be right?), but it appears as though you are ignoring irrefutably true observations. You just asserted that “the data” doesn’t show Claude introducing any bugs (which is a bizarre claim) after previously responding to a documented bug with a… deferral? Do bugs not count if you can find a vague excuse for it? There is nothing in the blog post that is evidence that Claude didn’t introduce bugs. It is a thought experiment that uses “increase bugs” and “increase bugs more than a given arbitrary statical amount that I selected” as interchangeable statements.
- themafia 4mo ago> If anyone complains about my verbosity or sentence structure — as they usually do, which is the reason I originally let the AI write the prose, among other reasons obsoleted by templating — they can go fuck themselves. You can write for an audience or you can write for yourself. Which is fine either way but you shouldn't pass the blame for bad results on to your audience. > and recieving almost no substantive input, discussion, or response on the actual content of the article Well did you write it for that purpose? > "Just wait, more bugs will surface" -- v3.4.3 has been out long enough Wait for _more releases_. As your own data shows the bug rate is not consistent between releases. So this is probably not a worthwhile metric. Perhaps systems touched, new features included, or attempted fixes would be a better way to contextualize releases and the goals of the author.
- lbrito 4mo agoWait, how is any of this relevant if there were only 2 Claude commits? My statistics courses are far behind me, but don't you need at least 30 data points to conclude anything?
- logicprog 4mo agoDepends on the methods you use. If you're trying to fit curves and so on, yes. The methods I use were designed for very low amounts of data, and are generally okay for that, specifically and especially when you're just trying to show a lack of evidence for some non-null hypothesis. And again, that's kind of the point. There's exactly zero actual evidence, however you slice it, that "Claude broke rsync" except cherry-picked anecdata, and the whole point of my analysis is to demonstrate the total lack of any such trend/evidence at all, and just how in-distribution/normal these releases are, to show that if people hadn't known Claude was involved in them, they wouldn't have remarked on them.
- wlonkly 4mo agoIt's not uncommon to have small amounts of data come out of experiments. These are appropriate tests for the size of the data. These tests failed to disprove the null hypothesis.
- matheusmoreira 4mo ago> My statistics courses are far behind me, but don't you need at least 30 data points to conclude anything? There is no fixed number. Sample size depends on the size of the set you're sampling, desired margin of error and confidence interval. If your total set has a million items, you need ~16600 samples to draw conclusions with 99% ±1% certainty.
- kelnos 4mo agoIt wasn't 2 Claude commits. It's 2 releases where the (many) commits were largely co-authored by Claude. > My statistics courses are far behind me, but don't you need at least 30 data points to conclude anything? That cuts both ways. If we say that the author here can't claim any conclusion because there are only 2 Claude-authored releases, then we must also say that the people claiming "Claude broke rsync" have no statistical basis to draw that conclusion, either.
- WesolyKubeczek 4mo agoThe discussions around this have devolved to excrement anyway, I feel tempted to invoke the meme where the goose asking a guy what his jacket is made of, asks “where is your reproducer case!?” instead. Instead we have a shitstorm over presumably legit issue, for which the only source is some mastodon post. One command that used to work in 3.4.1 and stopped working in 3.4.3. Just one! We could have already bisected the living shit out of this and go home, but no.
- steno132 4mo agoThis is just narrow thinking. Say Claude did increase the bugs in rsync by a negligible factor. So what? You've saved a significant amount of time for a decent number of humans, and if those humans are working on other projects, the overall net output for the world is net positive compared to without LLMs. You have to broaden your perspective. It's not just about how rsync was affected.
- boxed 4mo agoLet me translate this comment: > ok, so I was wrong and badly, but I will double down and say I was right anyway
- mmonaghan 4mo agoI think there's evolution at play here - if you dislike AI enough to opt out of using any ai-generated code, you will likely suffer. I think there's definitely a conversation to be had about whether to disclose AI use or not but that's a separate issue if you assume that everyone is using it in some respect.
- AEVL 4mo agoHow does the analysis look if we only count the >=90 severity cases—that is, if we downgrade the severity of all <90 cases to 0?
- logicprog 4mo agoFeel free to run it and find out. I don't think it would produce very much useful information though
- parliament32 4mo agoThank you for (re)writing this in your own voice. Despite how much effort might be put into methodology, data collection, etc.. reading slop is unbearable, full stop. It's not intentional, but I have almost a nauseated reaction when the "AI tone" comes though, regardless of how good the data or how accurate the writing is. Your verbosity and sentence structure are not a problem. I hope that publishing this gives you a bit more confidence in your writing, because it's legitimately good.
- tiahura 4mo agoWrite with your own voice and then polish with ai.
- dvt 4mo agoIt's always the most insufferable people that make the biggest hullabaloo about a project they have nothing to do with and have never contributed to. People with literally zero skin in the game using the AI boogeyman to push some agenda or some anti-agenda. OSS has become so incredibly toxic in the past decade, and consumers of OSS have become extremely entitled. I run a smallish project with ~1k stars and I've stopped maintaining it last year because people feel like they're absolutely owed features or bug-fixes or whatever. It's tiring and a complete shame that author has to make such an insane deep dive into a random accusation that just caught on social media. I want to emphasize that this has nothing to do with AI, it's just tech tourists, consumers (as opposed to creators), and engagement farmers that have taken over. AI slop probably doesn't help, but the underlying issue has been brewing for at least a decade. Also, the "making soup for the homeless & pissing in it" is not only an off-base analogy (software is pretty low on Maslow’s Hierarchy of Needs), but also somehow looks down on both people in need and the volunteers that help them. Just absolutely gross.
- Panino 4mo ago> It's always the most insufferable people that make the biggest hullabaloo about a project they have nothing to do with and have never contributed to. Agreed, and similarly, as a hobbyist programmer who loves Rust and Go, I've always felt that the people who command others to "rewrite it in xyz" are not themselves developers, they're "ideas people." There's a mass of these people whose main interactions with the world are through the dramatic forcing of their correct opinions. > I run a smallish project with ~1k stars and I've stopped maintaining it last year because people feel like they're absolutely owed features or bug-fixes or whatever. That's a bummer and it's something I'm fearful of. I post some code on my website, not on a github type site, and don't interact with people about it. It's nice and plenty of people do it. Is that something you'd consider?
- matheusmoreira 4mo agoAbsolutely agree. Quite a lot of judgement from people who benefited from this guy's software for over 20 years, probably without ever helping him pay his bills even once.
- 4mo ago
- WhereIsTheTruth 4mo agoLLMs don't create bugs, people do
- GodelNumbering 4mo agoWas just looking at commits and came across a commit and its revert original commit: https://github.com/RsyncProject/rsync/commit/d046525de39315d625ffaef4fdd6e7cf12148016 https://github.com/RsyncProject/rsync/commit/d046525de39315d... ``` - if (!ptr) - ptr = malloc(num * size); - else if (ptr == do_calloc) + if (!ptr || ptr == do_calloc) ptr = calloc(num, size); ``` Written with claude. This is a good example of what slips through LLM attention. It forces all allocations to be calloc as if it is a strict upgrade. For large and recursive allocations, this becomes a significant cost. reverted in https://github.com/RsyncProject/rsync/commit/7db73ad9a1b8721f14a43219d73127b23b86fe00 https://github.com/RsyncProject/rsync/commit/7db73ad9a1b8721... if you read the description of revert half carefully, it's easy to tell that even that was written by an LLM . I can understand the sentiment of whoever posted the original thread.
- wolletd 4mo agoAlso the amount of commits is suspicious. In the last two months, rsync had about as much commits as in the last two years before that. Most of them written with claude. And then stuff like this is in there. That's exactly what I'd expect when someone is excited about AI usage and becomes... well, sloppy.
- logicprog 4mo agoTridge already explains this: "Like many developers of open source packages I’ve been hit by a flood of security reports lately in my role as the rsync maintainer. Many of those reports are AI generated (not all though, there are some notable ones with very careful and high quality manual analysis). As this flood started to get more intense I realised I needed to raise the defences on rsync a lot — we needed much more thorough test suites, code coverage analysis, CI testing on a lot more platforms, deliberate and thorough scanning for possible security issues (so I find at least some of them before other people!) and the addition of a whole lot of defence-in-depth hardening techniques. This is all a huge amount of work. " https://medium.com/@tridge60/rsync-and-outrage-d9849599e5a0 https://medium.com/@tridge60/rsync-and-outrage-d9849599e5a0
- iainctduncan 4mo agoWhat strikes me about the post is that it goes to great lengths to talk about proper statistical methods, but then is written in the most clearly biased language ("what stupid AI haters get wrong etc). If you want people to take your study seriously, why wreck it by coming across with such a strong prior bias? I stopped reading...
- logicprog 4mo agoIf they're the statistical methods and metrics hold up, or they don't. Also, if you don't want to read my opinion on things, then just grab the GitHub repo and run the end-to-end replication and look at the output data yourself.
- int_19h 4mo agoTo be fair, the tone of the article is practically chill compared to the comments it is written in response to.
- jarym 4mo agoI've been coding for over 2 decades. I love it, I've always loved it and I likely always will. I was an AI skeptic some months ago but truly Claude and Codex have changed my development style and velocity in a way I never imagined would ever be possible. With that, yes, I produce more code and am finding more bugs. So looking over at comments in HN articles the amount of polarising hate to anything produced with AI is quite surprising. Just because some AI helped or even produced entirely doesn't suddenly make a project 'vibe coded' as if that's meant to be some insult levelled at users of LLMs. It reminds me a lot of when offshore outsources started getting more software development work from the mid-90s with all the derogatory remarks made towards 'Indian developers'. Now we're in the mid 2020s and similar remarks are made towards AI. I don't get it. I really don't. What I do know for sure is more and more code will be AI generated with or without the detractors.
- nomel 4mo agoI've always noticed, within any subject involving tools, there are people who like the tools, and some people who like to use the tools to do something else. With programming, I've always been in the later: it's a tool that allows me to do what I actually love, which is problem solving, system level thinking, and providing some nice solution to that problem, that happens to be through software. So, I have an absolute blast with AI, because it helps do the more boring bits. And, seeing my non-programming colleagues get excited to see their vibe coded ideas become reality has been so much fun. I'm genuinely curious to hear the perspective of someone anti-AI, who works in software. Perhaps the impending doom/skill shift of our profession?
- lelanthran 4mo ago> So, I have an absolute blast with AI, because it helps do the more boring bits. So... you're vibing? Not looking at the code at all?
- Joel_Mckay 4mo agoPersonally, it would still bother me if some lazy bro hit a code-generator and people end up dead. For context search, I find LLM quite useful... still wrong 20% of the time... but it has some utility. Here is a thought experiment: If "AI" will eventually generate your work, than what actual value do you bring to the table? =3
- RustyRussell 4mo agoFor those commenting, I suggest you read the post linked by the rsync author: https://medium.com/@tridge60/rsync-and-outrage-d9849599e5a0 https://medium.com/@tridge60/rsync-and-outrage-d9849599e5a0 (Disclosure: while I haven't talked with him in years, Tridge was my colleague and mentor for many years. I feel it is worth considering his view before joining a crusade)
- nullc 4mo agoI think that's an extremely well done response on his part.
- deleted 4mo ago[deleted]
- matheusmoreira 4mo agoThis should be the top comment. I think it's pretty sad that he even had to write it. Quite a lot of judgement from people who aren't paying his bills.
- dnnddidiej 4mo agoThe title at least sounds less like judgement and more analysis and more about AI assistance (and claude in particular) than rsync. Maybe I am too used to postmortems!
- el_io 4mo agoI think they're talking about the whole twitter and github issue things.
- Laurel1234 4mo agoYeah a big reason you see so much pushback on clanker slop is that it's having (and there was certainly the expectation of it having) a negative impact on the ability of plenty of people to pay their bills.
- cobertos 4mo agoThis post just gives me more questions than answers and I'm unable to form a decision: * Why was v3.4.1 the most buggy, right before the Claude commits? Why did "nobody notice"? It's way to strange to just say welp, it must be human error. * Why does v3.4.2 have 0 bugs, or 0 bug score. And why was such an outlier (no other commit seemingly has this??) allowed to mix into aggregate statistics and bring all the "is Claude buggy?" scores down. Tbh idk how that _wasn't_ a red flag in the author's analysis... This article feels like half of an analysis presented as a highly complex finished product due all the advanced stats they're running.
- logicprog 4mo ago> Why was v3.4.1 the most buggy, right before the Claude commits? Why did "nobody notice"? It's way to strange to just say welp, it must be human error. Why wouldn't it be except question begging priors assuming it couldn't be? > Why does v3.4.2 have 0 bugs, or 0 bug score. And why was such an outlier (no other commit seemingly has this??) allowed to mix into aggregate statistics and bring all the "is Claude buggy?" scores down. My original metrics which didn't filter out feature requests and questions had it at four bugs and prior to that it was even higher and it didn't make much of a difference to the overall analysis (fell well within the IQR, the lower end of it too). Also, removing one outlier just because it looks kind of funny to you, especially when we only have two Claude releases at all, would be worse in my opinion and more arbitrary.
- cobertos 4mo ago> Why wouldn't it be except question begging priors assuming it couldn't be? A multitude of reasons? A change in maintainer. A change in the mental state of a maintainer. A sudden focus by the community on a given undesirable behavior. Someone else here suggested use of Claude AI before it was disclosured. The framing implies that it was human-produced coding error, but my point is it could be _any other human error_ or even just some odd benign human behavior (a stampede of bug submitters), affecting the data. Which does not lead to the conclusion that AI code > human code. Not looking at these potentials is so unsatisfying. > My original metrics which didn't filter out feature requests... It still feels like a lot of weight of the phrase "If that doesn't look like a red flag to you, you'd be right." hinges on the fact that one of the versions has 0 bugs and it really killed the weight of that statement for me, because the oddity of there being 0 bugs just wasn't explained. --- Could you please post the duckdb file that has the raw bug -> severity + version mapping to the GitHub repo? I have a desire to dig into this myself
- gravypod 4mo agoThis is a really cool post but I think one metric we may want to also look at is does using agentic coding tools in one domain impact your coding abilities in another domain? A lot of people I know have been talking about getting rusty on the fundamentals recently. This is not something I am particularly feeling as I do a mix of running agents in parallel and writing some code manually where it makes sense. But if people who have been prompt-only at work come home and work on rsync and are more "rusty" maybe that could also lead to more bugs? This would be even harder to measure.
- 1a527dd5 4mo agoI'm amazed that this is still being discussed. It's open source, no one is forcing you to use it. If you don't trust the newer versions; use the old versions. If you no longer like the maintainer because of reasons, fork it/start your own. It's not that hard. Storm in a teacup.
- throw7 4mo agoTrust is slowly gained and easily lost. The amount of apologia I hear from top-tier developers signals an inflection point downward.
- MantisShrimp90 4mo agoI think this writer kinda took the bait which is fine someone had to do this so we couldn't debate endlessly. But the reality is that if you were already set enough to call rsync slop because of a single post, you aren't going to be more down now. Even in these responses I see everyone nitpicking and moving goalposts as if one more commit being actually claude-aided will tip the scales from stable project to "vibe coded slop". Software has always been fuzzy, we have never come up with an objective way to handle software quality, and this Uber hatred of llm contributions lets the humans who make egregious bugs and mistakes off the hook. Taking a step back, we need to have more empathy and thoughtfulness of one another in this space. Its new and people are experimenting and there will be nothing good coming from personal insults and DDOsing a good project just because someone got ragebaited on threads, x, mastodon or whatever else. How do we determine bugs and increase quality? Its almost like we have been grappling with this question for decades and I still hear people fight on the best way forward. Simple design, test driven development, user surveys, all of the above have been used as a proxy for software and they all failed to capture everything. Back in the day we used that ambiguity to give each other grace, now we use that ambiguity to tear down other creators. Whatever, if open source software really is dying its because of this toxic shit just as much as the llms
- thin_carapace 4mo ago'this toxic shit' would not be occurring if we didn't invent a machine that can be used either as a firehose or a scalpel. I do acknowledge that behaving hurtfully towards somebody giving something away for free is unwarranted behaviour. perhaps a universally agreed quality control method does not exist - this does not suggest that ai slop is anything but low quality code. ai can indeed be used well, however you yourself mentioned letting humans off the hook for making egregious mistakes. pushing out ai slop IS an egregious mistake. when a release contains more commits than the previous N releases, slop likelihood increases, therefore further evidence is required to prove non sloppiness.
- TZubiri 4mo agoI haven't used this thing for like 10 years, when my modus operandi was googling my question and installing whatever stackoverflow suggested. Can someone explain why one would ever use rsync (pre vibecode version) instead of cp and dd? Can't we just 'apt remove rsync' and save ourselves the time even spent on evaluating this dependency? Thanks
- int_19h 4mo agoBecause cp will copy everything, while rsync will copy only the things that actually need copying, and also delete the things that should be gone?
- Arcuru 4mo ago> rsync (remote sync) is a utility for transferring and synchronizing files between a computer and a storage drive and across networked computers by comparing the modification times and sizes of files. https://wikipedia.org/wiki/Rsync https://wikipedia.org/wiki/Rsync
- Joel_Mckay 4mo agoIf you deal with large numbers of files, the ability to dynamically skip compressing media and zipped files for transfer can be extremely handy. While stuff like sshfs is great for a few small files (and win11), it will be an order of magnitude slower than an rsync task. Most smart folks automate backup/recovery scripts, and only sometimes edit them with a new OS install. =3
- manlymuppet 4mo agoUnrelated, but this post has a level of rigor you rarely see nowadays. I think it deserves to be commended for that. HN relatively, is a very intellectual part of the internet, yet even still, it's really common to see very uneducated opinions here. Not that everyone needs to be very educated, but posts with plainly wrong assumptions and biases shouldn't go completely unchecked so rampantly.
- nelox 4mo agoThe peak of cascading effects from errant dependencies has yet to come
- aplomb1026 4mo ago[flagged]
- amluto 4mo agoReposting my previous comment because the post I commented on earlier was flagged to death: This is kind of a sad situation. Tridge is an excellect programmer and a very respected member of the community, and I totally get it. rsync, like most old C projects, has a lot of accumulated cruft, and things that would be nice to fix, and bugs. And those bugs come in at least three classes: semantic bugs, improper interactions with the OS, and memory safety bugs. And the author and long-time maintainer has the same problem as every other maintainer and team: not enough time to deal with everything. And now LLMs come along, and they are so, so seductive. They will fix your bugs if you ask them to. They will even find your bugs. And they're right a remarkably large fraction of the time. It's magic! You can write an agent loop or magic harness or swarm and let them do this on their own if you want. And so you start getting through your backlog, and it's fun, and you feel good, and you let your guard down. And you start having problems: - Your favorite LLM does not have the context that lives in your head. I use rsync because Tridge wrote a fine piece of software, and he knows how to write serious software, and I'm willing to accept that it's in C and therefore almost certainly has a safety bug or three. If I wanted to use claude-ersatz-rsync, I'd use that instead, but I really don't, TYVM. - Remember how LLMs are right a remarkable fraction of the time? The fraction is remarkable, but it's nowhere close to 100%. (Yet? Who knows. Right now, it's DEFINITELY nowhere near 100%.) - The training process for the current crop of LLMs does not adequately reinforce long-term maintainability of the outputs. And, for all the LLMs seem magic, they seem to love a workload in which they write code with poorly named functions and no docs and sort of assume that they can parse their own code down the road and figure out WTF is going on, and they are AT BEST only a tiny bit right. Because every project has interfaces where one module touches another, and every LLM has very limited context (larger than humans' in straight up verbatim working memory but MUCH MUCH WORSE than humans' (for now, anyway) in actual broad picture retention), and this workload doesn't work. If it did, we could give up on structured programming and just have the LLMs vomit up uncommented asm. And so, where humans have conventions and decently named functions and ideas that you shouldn't churn your code just for funsies (at least not in a production context), LLMs do this: https://github.com/RsyncProject/rsync/commit/30656c5e358b1c6 https://github.com/RsyncProject/rsync/commit/30656c5e358b1c6... Most of that is blindly changing calls do functions like do_foo(args) (which makes sense) to do_foo_at(the same args), which makes no sense. Sorry, but the world of POSIXish-targetting programers (including, presumably, Claude) knows what _at means, and it means "at" the specified directory fd. Which is not specified in the call sites. It makes no sense at all. Buried in all that mess [0] is the implementations, which are sloppy. Seriously: - There's a function called do_utimensat_at. Is Claude stuttering? - There's a lovely comment in syscall.c:1660-1673 that's quite bad. It's handling strings that contain "/../" and such. If there's some actual contract that the function makes to its callers (and there surely is -- this is critical security-sensitive code), then SAY WHAT THE CONTRACT IS. Don't bury a partial explanation in a comment in the middle. - There's a repeated pattern: In do_foobar_at(path), there is, in effect: if (!path) do_foobar(path); Nice NULL pointer handling. Is NULL a valid argument or not? Why handle it by forwarding it to the less secure variant? - Those nice, supposedly secure "at" variants check for paths that start with '/' and forward to the raw insecure syscall. And they don't check for .. in the middle. So what, exactly, is the special code for .. promising to do? (See above.) I don't think more details are needed. But my take is that this whole thing is a mistake. I personally work on the sort of code where messes like this are entirely unacceptable. And using an LLM while maintaining the kind of oversight that prevents it is mentally taxing and not exactly fun. If you want to fix all the gunk in a C program like rsync by LLM magic, go rewrite it in Rust or something -- you're already exposing yourself to a massive rewrite and all the risks that entails, and you're pretty much guaranteeing a high level of sloppiness, so at least use a language that is more resistant to slop. [0] Which GitHub doesn't even render by default because their diff viewer is so bad. [There were follow-ups. See https://news.ycombinator.com/item?id=48352182 https://news.ycombinator.com/item?id=48352182]
- igregoryca 4mo agoClaude in general probably increases observed bugs in rsync, because it can churn out vulnerability reports that necessitate tons of changes to software that people are accustomed to working flawlessly in non-pathological use cases. I don't have empirical evidence for this claim, but best I can tell, security patches are the principal source of observed bugs in software of a certain vintage, because they cause churn. (Just think of Windows updates that break drivers.)
- vlovich123 4mo agoIf the author is this concerned about security, I’m curious why rsync doesn’t just build with fil-c by default and skip the noise. Those who need the extra perf to do more than 1 gigabit/s can build it in “unsafe” mode.
- nilslindemann 4mo agoPlot twist: This blog post was written using Claude too.
- deleted 4mo ago[deleted]
- block_dagger 4mo agoDo people enjoy interrogative headlines? Find out at 11.
- xmddmx 4mo agoThere's a meta-level of irony here that's important to note. TFA is defending the use of AI, and it very clearly (to me) used AI to analyze the data and present the results. In doing so, the author used statistics in a way they do not appear to understand, and ended up making numerous false claims (you can see the thread discussing these here https://news.ycombinator.com/item?id=48417626 https://news.ycombinator.com/item?id=48417626 ) In short, the study doesn't have sufficient statistical power, and is making "no difference" claims that aren't justified. The meta-irony is this: the author used an LLM to interpret data in this study, and seems to have made the same category of mistake (confidently asserting falsehoods) that the study was supposed to be investigating (confidently submitting bad commits to the rsync project).
- classified 4mo agoAI is so much like a religion. There is nothing you can say to a believer that will make them question their believes. Or more generally, you cannot reason anyone out of something that they want to believe.
- Joel_Mckay 4mo agoIt gets pretty dark if you pull that thread of reason. =3 https://en.wikipedia.org/wiki/The_True_Believer https://en.wikipedia.org/wiki/The_True_Believer
- newsoftheday 4mo agoAI is nothing like religion. People behave similarly to AI when debating their favorite sports team, or for Java coders, Checked vs Runtime exceptions. Religion is about faith and what people feel and sense as much as believe.
- logicprog 4mo agoThe statistical methodology I used is mine. As is the interpretation. Completely. To the degree that I misunderstood statistics (and it is under debate even in the thread you link, and the people accusing me of misunderstanding statistics there are universally misrepresenting my point, which is to point out a total absence of evidence for any difference, not to prove the null hypothesis) that's on me
- eddysir 4mo ago[flagged]
- htk 4mo agoFlagged. Article is as AI heavy as the commits that people are complaining about.
- aarjaneiro 4mo ago> "Claude clearly made things worse" &emdash; the main claim Even this report is full of claude-introduced bugs
- runarberg 4mo agoFor those that don’t know the html entity is — not &emdash; although I think in modern codebases people usually just type — directly. This mistake does exist in the wild though: https://github.com/search?q=%26emdash%3B&type=code https://github.com/search?q=%26emdash%3B&type=code If I was more ambitious I would plot the dates of the blames of these results in a histogram and see if an there is a significant increase in these mistakes (over a baseline —) correlating with the release of some models.
- logicprog 4mo agoThat one was on me. I always mess that up.
- e40 4mo agoWhile I'm grateful for all Andrew has done to create and maintain rsync, I rely heavily on it for backing up files between machines on my home network, so I've spent the time to figure out how to pin the Homebrew version of rsync to 3.4.1 because the bugs in the subsequent two versions really scare me (as does the original report that triggered all this). Here is the process I used to do it, which was way more complex than I thought it would be: https://gist.github.com/e40/caa67c1b8d439a528695f996d0519d8e https://gist.github.com/e40/caa67c1b8d439a528695f996d0519d8e
- sathyayoshi 4mo ago[flagged]
- esailija 4mo agoWhat on earth is this. Literally the only thing that matters is are there more bugs after AI written code is allowed into the codebase at all. We all know the answer to that lol. But it's always nice to see "data" can be used to make any conclusion you need.
- logicprog 4mo agoThe data literally shows there aren't, there have been worse releases before. In what way did I manipulate the data?
- nasretdinov 4mo agoRegardless of the claims made in this analysis, I've personally observed that there are indeed more bugs (or more subtle issues, like nonsensical error messages) being shipped when using LLMs, but not _really_ because LLMs suck, but because you're spending less time thinking about the problem, and you subsequently miss more edge cases, etc. The best approach I've tried that actually increases quality (and _may_ speed up development) is to write ~80% of the code yourself and then ask LLM to review it thoroughly. While it's doing its thing you're also thinking about the code and reviewing it yourself in parallel. You then merge the findings and fix stuff worth fixing. At this point the authorship of the code is still mostly yours, you _understand_ the system and you ship fewer bugs, slightly faster than otherwise. It's a moderate improvement to the workflow, but it actually doesn't cost nearly as much either, and definitely doesn't produce rage at the machine from the slop. The only downside is that it requires lots of discipline, and it's a relatively rare commodity among software engineers these days.
- Havoc 4mo agoWhat’s the deal with anti ai people being so rude
- rswail 4mo agoPersonally, I'm going to believe tridge, someone that has contributed more to software than 99.9(recurring) of the software development community over the last 3+ decades, than a bunch of brigaders jumping on the anti-AI backlash. There was one regression bug apparently (related to multiple destinations and the way people do backups), but all the attention/anger has been about a test suite that makes the rsync development better, more rigorous and copes with the onslaught of both good and bad AI generated PRs as well as hardening something that has two decades of C code in it. People need to grow up and appreciate what others in the community (especially people like tridge) have provided.
- nazgul17 4mo agoIn a scenario like this, where we can only know if the code has bugs, not if it doesn't, isn't survival analysis a more appropriate statistical technique? I.e. a technique where time is a first-class citizen. By the way, I did find this a bit hard to read but, as instructed by OP, I'll go fuck myself. For what it's worth, I find AI written prose easy to read, and am annoyed by all the constant HN comments which just point out the author was AI, without anything else substantive to add.
- foxes 4mo agoI wonder if all the commits which involve adding tons more test are the basis for a rewrite in rust anthropic marketing event
- deleted 4mo ago[deleted]
- ltbarcly3 4mo agoI'm noticing more and more AI writing everywhere, from youtube to this article: From the subtitle: "Nothing complicated, answers only one question: " clearly LLM generated.
- ch_fr 4mo agoThis article is a rant disguised as data analysis. I don't know how to word this in a non-confrontational, respectful way, but this article just feels like ammo for your next "debate with your anti-AI ennemies" where you get to say "look, I proved with data that those people had a disproportionate reaction and have double standards, therefore anyone who dislikes LLMs or their impact are the same!". Like, sorry, I know that sounds really reductive, but this really is the vibe I get when reading this and your other replies where you repeatedly talk about "showing the hypocrisy and double standards". The global LLM discourse has grown massive, it spans trillions of dollars in promises and investments and affects pretty much everyone, so it's the easiest thing in the world for both sides to just find some people being assholes in the other camp and say "look, here's how [other camp] behaves". The irrational, extreme, and heinous reactions are partly bandwagoning, and you can go about your day thinking that anyone who reacts like that is evil. But if you wanna dig a bit further, you'll notice that the entire media sphere has been screaming in everyone's ears for a few years now, that they're expandable, low-value-human-capital. All the money in the world (exaggerating a little) is being spent on making sure to remind anyone who opens a computer, opens a website, looks at a billboard, or turns on his tv... that their boss really really really wants to replace them. Now you'll say that the friendly rsync contributor has nothing to do with any of this and... well yeah he doesn't. You don't need to agree with an emotional response to understand where it's coming from, and even if you're still dead set on considering them "the enemy", then understanding why the anti-AI crowd reacts like that is STILL a positive for you.
- pie_flavor 4mo agoWhy is the guy being rigorous worthy of criticism, but the guys being idiots aren't? Did you post any similar calm-down comments in either of the HN threads on the original attacks?
- ch_fr 4mo agoI am more inclined to be critical of AI boosters, so what? Am I supposed to crumble under the weight of immense cognitive dissonance because I have... a stance in the discourse? These guys on the github thread aren't my friends, I have no concern for them embarrassing themselves or leaving a bad digital footprint by drawing ms paint gore. I also have no concern for OP, but it just so happened to be the post I found, and I just so happened to be in the mood to leave a comment. Engaging in LLM discourse is already a waste of my time, I'm not going to waste more of it just to avoid fallacious accusations of double standards because I didn't "do the same for the other side".
- bwfan123 4mo agoSoftware is only as healthy as that of the mental-models of its "human" maintainers. This was implicit in the writings of Peter Naur (programming as theory building) and the Fred Brooks (mythical man month) ages ago. AI as a tool can assist just as IDEs and linters have assisted. But eventually the mental models of the human maintainers is the gate and bottleneck. Applying this to rsync, its maintainer would need to foster and grow other human contributors to eventually become maintainers so that the human mental models are carried forward.
- moktonar 4mo agoI think we should start having ai versions like beta ora alpha versions and then consolidate them into human made versions with time, at least one is free to stay safe or on the bleeding edge as one likes and we all get a win-win best-of-all-worlds situation (hopefully)
- ladax72707 4mo agoSo a project is using a GPL licence, but instead of forking you harass the authors and you somehow think that you are the smart one and that you are doing anyone a favour?
- guilhas 4mo agoRsync is a highly trusted software, included in many distros. To move important, and high quantity If several or critical lines of code get changes quickly, and keeps breaking things, with or without llms, there will be backlash Rsync should rightly loose reputation if the project allows the release breaking changes to follow the latest hype trend
- drankinatty 4mo agoI've got no love for AI, don't use it, but also after writing code for more than 40 years, keeping things in perspective helps. Whether it's you pecking away or some coding assistant helping, there will always be the potential for a regression or two to fly under the radar. (not like I've never done that before... nope.) The issue the coding tools like Claude present is the sheer size and scope of changes and commits they generate that would take mere mortals months of careful coding to do. That's an issue everyone using those tools will have to confront. I don't know Andrew personally, from a "let's go have a beer" standpoint, but I've known him from the samba list and his work with rsync for a very long time. My take on the issue is less about the regressions and Claude screw-ups and more the lesson to all about the reliability of the coding tools and the diligence required to validate what they spit out. It's an unfortunate black-eye, no doubt, but it's not a unique one. The takeaway is if something like this can slip by somebody like Andrew, then we all need to redouble the validation effort, lest we too are destined to share an unfortunate black-eye or two. Never forget, "to err is human, but to really foul things up requires a computer." AI just applies that adage at industrial-scale.
- sspoisk 4mo ago[flagged]
- zzo38computer 4mo agoI agree that the bug report is not very good. I also agree that having commits written by Claude is not necessarily what caused the bug, although it might be (it is also possible that some of them introduced bugs and others didn't); whether or not it is in this case, is I don't know (some people think it is, but some think not). (Software without code written by generative AI will still have bugs too.) However, the claim that "the original post was [...] no bug report" seems wrong; it does have a bug report, although not a very good one. It says that incremental backups using multiple --compare-dest arguments do not work, so it is a bug report. But, it should have been written differently, including by putting the text directly instead of a screenshot, giving a proper title, better details about the bug being reported, etc. Their claims that they introduced deliberate bugs, are unlikely to be accurate, and not worth making those claims nor the violence that they involve. I do have reasons for not wanting LLMs to commit code, so I agree with their opinion about that, but that does not justify making a bad bug report and the other stuff that they did. If it is FOSS, someone who disagrees with the project can fork it and make their own version, as has been done with other FOSS projects as well. I think it is good that they are making statistical analysis. However, they used a language model to classify bug reports. They mentioned some things that might be missed, and they could be missed whether or not you are using a language model to classify bug reports, although there are some other possibilities e.g. whether or not a single report should count as multiple bugs in some cases, and mistakes in marking reports as duplicate.