9 ms·
Conservation of Intent: why A/B tests aren’t as effective as they look
- a-dub 8y agoFrustration is not linear. Film at 11.
- birken 8y agoI could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed the 2 cohorts of customers and watched their behavior. The control group vs the 10% more from whatever the test was that increased conversion. They behaved exactly the same. They came back again at the same rates. They made the same amount of profit (per customer). Their response rates to emails were the same. They closed jobs at the same rates. As far as we could tell they were identical. I have to admit I was a little surprised too, but for our business it didn't seem this "high-intent" vs "low-intent" distinction existed. And with that out of the way we continued to optimize conversion rates, and our revenue continued to go up. Every company is different so I don't want to generalize too much, but if somebody tells me they ran an A/B test that said some key flow went up 10%, but then afterwards the traffic/revenue/whatever didn't go up 10%, I think the most likely candidate is bad test design. Humans are really good at rigging A/B tests to produce wrong results in their favor. I guarantee every single company who isn't maniacal about A/B testing does at least one of the following: - Uses a tool to grade A/B tests that isn't statistically sound - Let's people check tests too often and allows them to stop the test when it hits a good result - Running a test with a lot of similar variations and cherry picking the best one - Doesn't plan for enough traffic to detect the percentage of change their test is likely to produce All of these create the potential for the perceived gains of the A/B test not matching up with real world result. I'm not saying the distinction between "low-intent" and "high-intent" customers doesn't exist, but it is fairly easy to test for. Do that test for your business and see if that distinction exists. But don't use it as some magical explanation for why your A/B tests aren't producing the results you want as this article suggests.
- sokoloff 8y ago> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.
- alexbeloi 8y agoThat is actually one of the biggest contributing factors to the replication crisis in science, lots of scientists have been making this error for decades. Very not obvious.
- birken 8y agoI completely agree. We went through all of this stuff as our A/B testing was evolving at Thumbtack. They seem so simple, but the deeper you get you realize they are only simple if you don't "cheat", which everybody does. And you'll only realize you are cheating in the first place if you have somebody who is maniacal about testing. If you create a culture where positive A/B tests are lauded (which is good!), then you create a lot of people who want A/B tests to finish in positive ways. For those people, it doesn't really matter if their A/B test actually improves things, only if it looks like it does. This isn't nefarious, this is just human nature, but it creates a lot of creativity and energy at finding ways to making winning A/B tests. We'd have people run a test, where their new variation got off to a bad start, then say "oh, it was a bug", then they'd fix some irrelevant thing and start over just to reset the counters. I was guilty of it sometimes. You get excited about your tests and want them to win. That is why it is critical either to have gatekeepers like an analytics team to keep you honest or have a really specific protocol on how your company runs tests and only consider results of tests that followed the protocol.
- marcosdumay 8y agoWell, you can arrange the test so checking often is not a problem. It may even be the optimal way. But yes, statistics is not obvious at all.
- 8y ago
- jbob2000 8y agoWait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked. A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.
- matt4077 8y ago....and criticising statistics is the laziest kind of scientific criticism.... A/B tests are fine. They work. They allow inference of causality. They are easy to understand, and can be fun to run. They get you 90% of wherever you want to go, and such over-the-top criticism just seems like badly executed pretentiousness.
- jbob2000 8y agoYou’re partly right - a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important. But to the article’s point, if you’re running banking software or something, your users don’t give a shit what the button colours are; they will slog through whatever you develop because they need to get stuff done. A/B tests are a small tool that sometimes get taken too far or used in the wrong context.
- ddebernardy 8y ago> a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important. Not saying A/B testing shades of blue like Google reportedly does is anything useful, but picking a different button color altogether reportedly can make a difference. Even if it's 1% more sign-ups each month, that can translate to a few more potential clients per month - and more word of mouth. If you think that is too trivial a difference to matter, think about what 1% more interest would mean over your career as you compound interest for your retirement.
- lostcolony 8y ago'Taken too far or used in the wrong context' - sure. In the language of this article, banking app users are mostly all 'high intent'. But that doesn't mean you can't evaluate criteria other than users who completed the workflow to determine what is a design improvement. You can still measure time to completion, how long it took the user from entering the workflow to completing a given task as a measure, and play with the design. Optimize the things users are doing the most, that sort of thing. A/B testing can help you there. It's not the be all end all; you still need sound UX design to figure out what designs to test out, but it can give you measurable data as to what works, rather than just UX gut feeling, or purely lab based results which don't reflect reality.
- gfodor 8y agoThe title is misleading relative to the article's content. Surely, as the author points out, sometimes A/B tests can be misleading especially if you ignore longer term cohort analysis, etc. But often times, if you fix an obviously broken part of your funnel, particularly in the early acquisition stages, you're fixing things that are universally lifting the amount of people who ultimately are able to engage with your brand and product to the point where they can even form intent. The reality is most people are only willing to give you a tiny bit of their time during their first one or two engagements with your brand, so at that stage you're trying to sell them on your product, and build intent. A/B testing helps reduce the friction needed to get them through the core of your sales pitch. It's easy to come up with a thought experiment that shows A/B testing can sometimes be as simple as you'd imagine: just break the site. Your conversion drops to 0%, now split test the fix. Like magic, your control stays at 0% and your variant returns to normal. Nothing about "intent" in this scenario, this is pure friction resolution. Just a thought experiment, but shows that surely there are plenty of places where pure A/B testing and removing friction is a net positive without any fretting over this "conservation of intent" issue.
- foobaw 8y agoSlightly relevant but useful: use meditation modeling (https://eng.uber.com/mediation-modeling/ https://eng.uber.com/mediation-modeling/)
- mwexler 8y agoThis is a nice read. BTW, in case folks get confused, moderators are different from mediators. http://psych.wisc.edu/henriques/mediator.html http://psych.wisc.edu/henriques/mediator.html
- ben509 8y agoI think the idea of "high intent" is the same fallacy as the notion of "affordable." We say something is "affordable" because we have "enough" money to buy it, but that's not how people make decisions in aggregate. The reason economists talk about opportunity cost is because people are constantly optimizing decisions based on new information. (Humans may not deal with prices and numbers very well, but they're pretty well evolved to break time into chunks and work out plans to solve problems.) If you talk to an individual, they might say "I can't afford it," or you may talk to someone who didn't click through and they might say, "I was just browsing." The fallacy behind both is you're creating archetypes and assuming they represent the modes of the population. And even if you talk to the individuals you based those archetypes on, there is a whole history behind how they arrived at "I can't afford it." Those changing circumstances are why the aggregate behavior doesn't show some arbitrary level of "affordability," and instead you see a smooth curve of consumer demand. And the opportunity cost of continuing to view a web page will not have neatly quantized levels of intent, but rather individuals have a broad array of competing interests.
- snovv_crash 8y agoA/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even really visit anymore. Yes, I'm guity myself of clicking on those articles when I do visit, but at a certain point I've found that Ars doesn't have the news I'm after, so I turn elsewhere and now Ars has one less viewer.
- smueller1234 8y agoI think what you say is practically spot on. Yet I'd like to add that I don't think that testing frameworks (at this point it would be misleading to call them strictly A/B testing) HAVE to only reflect short term gains. It's hard to come up with (proxy) metrics that hold that kind of short term optimization in check. Just to pick one fairly obvious example for e-commerce: you can track returns/customer service contacts as a health metric or even explicitly assign a value to them to fold them into the primary metric you're optimizing. These kinds of safeguards typically need longer recording periods, so it's potentially a lot of data science and engineering effort to build tools that can handle the long term data collection and analysis. But it's not impossible. It's rather something people love to pretend is not a problem.
- vintagedave 8y ago> I've seen this repeatedly with ArsTechnica Can you expand? I am an Ars reader, and I too find it frustrating compared to what it used to be. I wish there were more in-depth technical articles; I find it too light on details, written for a non-tech audience. I really, really wish it had solid technical content, since it's what got me reading it. But I don't characterise it as click-bait (it seems clear) nor as especially political, except in so far as politics intersects technology and climate. In both those it seems balanced or erring towards freedom.
- kolpa 8y agoPolitics and a Climate and Freedom are interesting, but not what makes Ars Technica interesting.
- smueller1234 8y agoWhat the article largely discusses seems to be a problem with the metrics one chooses as a proxy. Let's say you're actually trying to optimize total transaction value on the site or total number of transactions or something like the overall fraction of users with at least one transaction within a certain window of time. Then - as the article rightly observes - getting users not to bounce on a particular page is a TERRIBLE proxy to what you're optimizing for. If that's not clear to you, you have no business running A/B tests without supervision. Source: co-designed one iteration of the experimentation framework for Booking.com many years ago. Indirectly managed the team of much more qualified people that took it a world further.
- User23 8y agoOne of my coworkers is a trained particle physicist and he informs me he almost never sees properly designed experiments used by our A/B testers. The result is that the testers almost always find what they are looking to find.
- MaxBarraclough 8y ago> You ship an experiment that’s +10% in your conversion funnel. Then your revenue/installs/whatever goes up by +10% right? Wrong :( Turns out usually it goes up a little bit, or maybe not at all. Never mind "The difference between high- and low-intent users", this could be explained in terms of regression-toward-the-mean, a phenomenon mention in neither the article, nor the discussion here. Have 1000 students do an IQ test. Pick the top 20 students. Have them do another IQ test next week. Their mean score second time round will almost certainly be lower than their mean score first time round. The reason they made the top 20 the first time round was a combination of having a high true IQ, and being lucky on the day. Second time round, they aren't 'defined to be lucky', as it were. It's the reason movie sequels tend to be worse than the original. The reason the sequel was made was that the original movie was far more successful than the average movie, on account of both unusually skillful creators, and unusually good luck. Second time round, you can't count on the luck component again.
- baybal2 8y agoTotally true, I list count of people coming from the web/startupey scene A/B testing their companies/business units into insolvency