6 ms·
The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. I don't believe A/B testing is an inherent evil, you need
by krisbolton 7mo ago
The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. That being said, vastly reducing an LLMs effectiveness as part of an A/B test isn't acceptable which appears to be the case here.
- ramoz 7mo agoI apologize for doing this - and I agree. I will revise
- s3p 7mo agoI still think you have a point here. Doing this kind of testing on users unwittingly is unethical in my opinion
- tomalbrc 7mo agoWould love to know why you would consider invoking Meta “a little much”. Sounds more than appropriate.
- krisbolton 7mo agoNot to start an internet argument -- I don't think it is appropriate in this context. A/B testing the features of a web app is not unexpected or unethical. So invoking the memory of cambridge analytica (etc) is disproportionate. It's far more legitimate to just discuss how much A/B testing should negatively affect a user. I don't have an answer and it's an interesting and relevant question.
- mschuster91 7mo ago> A/B testing the features of a web app is not unexpected or unethical. It's not "unexpected" but it is still unethical. In ye olde days, you had something like "release notes" with software, and you could inform yourself what changed instead of having to question your memory "didn't there exist a button just yesterday?" all the time. Or you could simply refuse to install the update, or you could run acceptance tests and raise flags with the vendor if your acceptance tests caused issues with your workflow. Now with everything and their dog turning SaaS for that sweet sweet recurring revenue and people jerking themselves off over "rapid deployment", with the one doing the most deployments a day winning the contest? Dozens if not hundreds of "releases" a day, and in the worst case, you learn the new workflow only for it to be reverted without notice again. Or half your users get the A bucket, the other half gets the B bucket, and a few users get the C bucket, so no one can answer issues that users in the other bucket have. Gaslighting on a million people scale. It sucks and I wish everyone doing this only debilitating pain in their life. Just a bit of revenge for all the pain you caused to your users in the endless pursuit for 0.0001% more growth.
- xg15 7mo ago> It's far more legitimate to just discuss how much A/B testing should negatively affect a user. I don't have an answer and it's an interesting and relevant question. You don't have an answer on "how much should A/B testing negatively affect a user"? So "a lot" would be on the table?
- SlinkyOnStairs 7mo ago> I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. I disagree in the case of LLMs. AI already has a massive problem in reproducibility and reliability, and AI firms gleefully kick this problem down to the users. "Never trust it's output". It's already enough of a pain in the ass to constrain these systems without the companies silently changing things around. And this also pretty much ruins any attempt to research Claude Code's long term effectiveness in an organisation. Any negative result can now be thrown straight into the trash because of the chance Anthropic put you on the wrong side of an A/B test. > That being said, vastly reducing an LLMs effectiveness as part of an A/B test isn't acceptable which appears to be the case here. The open question here is whether or not they were doing similar things to their other products. Claude Code shitting out a bad function is annoying but should be caught in review. People use LLMs for things like hiring. An undeclared A-B test there would be ethically horrendous and a legal nightmare for the client.
- airza 7mo agoIsn’t the horrendous ethical and legal decision delegating your hiring process to a black box?
- vova_hn2 7mo ago> ethical and legal decision These are two very different things. I suspect that in some cases pointing finger at a black box instead of actually explaining your decisions can actually shield you from legal liability...
- paulryanrogers 7mo agoFor some proponents, AI is liability washing
- garciasn 7mo ago> And this also pretty much ruins any attempt to research Claude Code's long term effectiveness in an organisation. Any negative result can now be thrown straight into the trash because of the chance Anthropic put you on the wrong side of an A/B test. LLMs are non-deterministic anyway, as you note above with your comment on the 'reproducibility' issue. So; any sort of research into CC's long-term effectiveness would already have taken into account that you can run it 15x in a row and get a different response every time.
- everdrive 7mo ago>I don't believe A/B testing is an inherent evil, Evil might be a stretch, but I really hate A/B testing. Some feature or UI component you relied on is now different, with no warning, and you ask a coworker about it, and they have no idea what you're talking about. Usually, the change is for the worse, but gets implemented anyway. I'm sure the teams responsible have "objective" "data" which "proves" it's the right direction, but the reality of it is often the opposite.
- cosmic_cheese 7mo ago> I'm sure the teams responsible have "objective" "data" which "proves" it's the right direction, but the reality of it is often the opposite. In my experience all manner of analytics data frequently gets misused to support whatever narrative the product manager wants it to support. With enough massaging you can make “objective” numbers say anything, especially if you do underhanded things like bury a previously popular feature three modals deep or put it behind a flag. “Oh would you look at that, nobody uses this feature any more! Must be safe to remove it.”
- deleted 7mo ago[deleted]
- mschuster91 7mo ago> The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. No. Users aren't free test guinea pigs. A/B testing cannot be done ethically unless you actively point out to users that they are being A/B tested and offering the users a way to opt out, but that in turn ruins a large part of the promise behind A/B tests.
- bcrl 7mo agoPlease name a computer science program that has an ethics component. Yes, I wish software developers were more like actual engineers in this regard.
- gnabgib 7mo agoAll Computer Engineering & Systems Engineering programs in Canada require two ethics components (once at graduation, once at P.Eng)
- ryandrake 7mo agoSadly, in the USA, I believe most engineering ethics classes are optional electives, and it shows when you look at the graduating student body today.
- saltcured 7mo agoYeah, and if you don't already have an IRB, your organization probably isn't ready to be doing such things responsibly...
- mschuster91 7mo agoMeta has had an IRB for well over a decade (following a scandal where they used their users as lab rats) and that didn't stop them from doing any of the BS they did ever since.
- hollow-moe 7mo agoTech companies really have issues with "informed and conscious consent" doesn't they
- cyanydeez 7mo agoRelying on a paid service for anything significant is basically accepting the Company Store feudal serfdom. Enshittification is coming for AI.
- xg15 7mo ago> The framing of A/B testing as a "silent experimentation on users" Sorry, but how is A/B testing not exactly that? The experiments may be on non-disruptive things like button color, but they're experiments no less. The users are also rarely informed about the experiment taking place, let alone on the motivation or evaluation criteria.