6 ms·
Here they mention that each bisect ran a large number of times to try and catch the rare failure. Reminds me of a previous experience: We had a large integrat
by Laremere 3y ago
Here they mention that each bisect ran a large number of times to try and catch the rare failure. Reminds me of a previous experience:
We had a large integration test suite. It made calls to an external service, and took ~45 minutes to fully run. Since it needed an exclusive lock on an external account, it could only run a few tests at a time. We started getting random failures, so we were in a tough spot: bisecting didn't work because the failure wasn't consistent, and you couldn't run a single version of a test enough times to verify that a given version definitely did or didn't have the failure in any practical way. I ended up triggering a spread of runs over night, and then used Bayesian statistics to hone in on where the failure was introduced. I felt mighty proud about figuring that out.
Unfortunately, it turns out the tests were more likely to pass at night when the systems were under less strain, so my prior for the failure rate was off and all the math afterwards pointed to the wrong range of commits.
Ultimately, the breakage got worse and I just read through a large number of changes trying to find a likely culprit. After finally finding the change, I went to fix it only to see that the breakage had been fixed by a different team a hour or so before. It turned out to be one of our dependencies turning on a feature by slowly increasing the probability it was used. So when the feature was on it broke our tests.
- n49o7 3y agoProbabilistic feature flags! Love it.
- Thorrez 3y agoAlways base the probability on something stable, such as hash of the username.
- deleted 3y ago[deleted]
- IshKebab 3y agoBug report: changing my username breaks $product. Yeah no thanks. It's probably better than completely random but software should be predictable and unsurprising.
- burnished 3y agoThe important part is the stability - if your usernames can change then they aren't stable so you don't select it. I think it is a good reminder that most things you think of as being unchanging that are also directly related to a person.. aren't unchanging. Or at least any conceivable attribute probably has some compelling reason why some one will need to change it.
- robocat 3y ago> changing my username breaks $product. https://m.youtube.com/watch?v=r-TLSBdHe1A&t=14m10s https://m.youtube.com/watch?v=r-TLSBdHe1A&t=14m10s Discussing a performance regression due to longer username due to username being in ENVIRONMENT variable which changes memory layout of process.
- btilly 3y agoI've used the hash of username+string trick before for a flag. I used it to replace a home-grown heavyweight A/B testing framework which had turned into a performance bottleneck. It worked quite well.
- dietr1ch 3y agoThat's why you have internal user ids instead of using data directly provided by users. Will it cost an extra lookup? It's cheap, and if you really need to, you could embed the lookup in some encrypted cookie so you can verify you approved some name->id mapping recently without doing a lookup.
- hedora 3y agoWait, we're talking about maliciously injecting bugs into your employer's software so they have the maximum impact, right? Clearly, making sure that 1% of all teams gets fired for being unable to run unit tests, then slowly ramping that by a few percent each review cycle is a good strategy. Ideally, the probability of breaking would drop off exponentially as you moved up the org chart. Something like "p ^ 1/hops_to_director_of_engineering" would work well. The trick would be getting the dependency to query ldap without being detected...
- Thorrez 3y ago
- patmorgan23 3y agoYou just need to be able to look up what feature flags where enabled on a given request (maybe by correlation I'd)
- robertlagrant 3y agoOr have the username be a number that is all the feature flags when converted to a binary representation. Then you can just have one username for each combination you want to test.
- gurchik 3y agoYou just need to make sure that this doesn't mean people are consistently "lucky" or "unlucky." I was on a team where app updates were deployed using a canary system. A small percentage of users (say, 1%) received the update first, then the team watched for incoming crash reports from that cohort. If it looked good, the feature was rolled out to a few more people, and this was repeated. This allows you to identify a problem by only negatively impacting a relatively small percentage of customers. The problem occurs when the calculation to determine which cohort the user belongs to is deterministic. In this case, the calculation was based on the internal ID of the user. This means some users always get the updates first, and deal with bugs more frequently than other users. Conversely, some users are so high in the list that they virtually never get an update until it's been tested by a wide user base, so their experience is consistently stable. Or you might have a problem where some players in a video game consistently take more damage than their friends: https://news.ycombinator.com/item?id=34742505 https://news.ycombinator.com/item?id=34742505
- thehappypm 3y agoMulti-armed bandits utilize this
- ambicapter 3y ago> Ultimately, it turned out to be one of our dependencies turning on a feature by slowly increasing the probability it was used. Wow. I feel like this dependency should be named and shamed.
- Laremere 3y agoBig company internal dependency. So nothing for the public to care about.
- vamega 3y agoWhat company. I've seen this being done (and my team does it a lot at Amazon) but curious to know if others are doing it at build time too. If done in a company with a monorepo I'd be especially interested in hearing more
- aeyes 3y ago> If done in a company with a monorepo I'd be especially interested in hearing more Are there any big companies left which haven't adopted a monorepo?
- deleted 3y ago[deleted]
- PartiallyTyped 3y agoAWS. We probably have the worst build systems :(
- scubbo 3y agoThis is astonishing. The build (and deploy) systems are, by a considerable margin, the things I miss most about having left Amazon (CDO, not AWS, but still). What do you dislike about them?
- yojo 3y agoYikes! FWIW, I think best practice here is to hardcode all feature flags to off in the integration test suite, unless explicitly overwritten in a test. Otherwise you risk exactly these sorts of heisenbugs. At a BigCo that’s probably going to require coordinating with an internal tools team, but worth getting it on their backlog. All tests should be as deterministic as possible, and this goes double for integration tests that can flake for reasons outside of the code.
- nosefrog 3y agoBut then you won't catch the bug before it hits production :)
- dmoy 3y agoAlso you end up with some strange long term test behavior. Because people will often leave feature flags in place long after full release (years sometimes), you end up with a default-off-in-tests only testing behavior with everything newer than N years since the last feature flag cleanup disabled. Yes it's kinda fractal of bad practices that have to align for this problem to occur, but that's the nature of tech debt.
- yojo 3y agoI agree that this is a real and separate problem, but I believe the solution lies outside of the test suite. One way I have seen this handled is to enforce restricting rollouts of a feature flag to 95% at most. That way turning a feature all the way on requires removing the flag from your codebase. It’s draconian, but honestly anything less than that leads to the situation you describe.
- dmoy 3y agoI like that idea a lot. We've been informally doing it on my current team, made easier since we can sort of cleanly do atomic code+flag updates in a single commit
- 3y ago
- painted-now 3y agoMan, this story sounds like you could be on my team :-) Pretty much experienced the same stuff working at BigCo! In the end, I think the real problem is that you can't test all combinations of experiments. I don't trust "all off" or "all on" testing. In my book, you should indeed sample from the true distribution of experiments that real users see. Yes, you get flaky tests, but you also actually test what matters most, i.e. what users will - statistically - see.
- joosters 3y agoThis sounds like a situation that would benefit from using an approach like all-pairs testing - https://en.wikipedia.org/wiki/All-pairs_testing https://en.wikipedia.org/wiki/All-pairs_testing Basically, if you have N different features (let's assume they are all on/off switches, but it works for multi-values too), in theory you'd need to run 2^N tests to cover them all, which would become completely impractical. But, you can generate a far, far smaller set of test setups that guarantee that every pair of features gets tested together. Run those tests and you'll probably encounter most feature-interaction bugs in a much quicker time.
- cscheid 3y agoAll-pairs is for _pairs of features_. For subsets you're in much deeper trouble because of the exponential dependence on N. For a fixed polynomial dependence, you can get clever and let tail bounds eventually work for you, but for exponentially growing hypothesis sets, that won't work.
- Noumenon72 3y agoI don't understand what you're saying. That it doesn't work if you need combinations of 3 or more features (a subset?)
- cscheid 3y agoThis is relatively subtle stuff, but here's an attempt at describing the general problem. I'm going to describe the deterministic case, but the probabilistic case is effectively the same. Let's say you have a bug you suspect is from an interaction of any one pair of 10 features being "on" or "off", but you don't know which specific pair causes the problem. Encode each of the states you could set up your code by a 10-digit binary string: 0000000000, 0000000001, 0000000010, 0000000011, etc. We could try the 45 possibilities in some order, and we would expect that on average it'd take us 22.5 tries to find the bug. But notice how your "target set" is smaller than the universe of strings: there's only 45 pairs of features, but 1024 strings. What happens if we try a random string of ones and zeros? Now, instead of catching just one possible pair, we are covering many pairs. The only problem is that we now won't be able to know exactly which pair caused the problem when it does. But we can build a corpus of strings that don't trigger the error vs. strings that trigger the error, and a random sampling soon converges on the correct pair. If you think about why this works, it's because any of these random strings has about a 1/4 chance to trigger the bug: wlog we can reorder the bits so that the buggy feature are the first two digits, and then we see that we have a 1/4 chance of hitting "11" on those two digits. The problem is that as you increase the size of the subset that needs to be active, the probability that your random strings will actually catch the bug decreases exponentially. For any _fixed_ target size k (the number of features that need to be active), the overall complexity is still polynomial in n (the number of existing features). But if k is a constant fraction of n, then this technique takes exponential time in n.
- ASinclair 3y agoThis is my daily life at BigCo. These bugs are the worst.
- tus666 3y agoEh this is what version pinning is for. Using edge can always lead to random breakagr, feature flags or not.
- sowbug 3y agoThat depends whether the development model is branch-stable or trunk-stable.
- hedora 3y agoThat wouldn't work here though, the dependency started by breaking an almost-undetectable fraction of the time. Imagine a scenario where your upstream dependency started out with one failure per 1,000,000 machine hours, then removed a zero once every 12 months. If you had 100 machines running tests at 100% efficiency, the bug would hit about once a year for the first year, then 10x the next year, and so on. Put another way, if upstream is malicious, and you're not auditing every line of their source code, you're screwed.