5 ms·
I generally agree, but there's one thing that bothers me about the practice: Before disabling the test, does anyone even consider the possibility that maybe it
by quicklime 6y ago
I generally agree, but there's one thing that bothers me about the practice:
Before disabling the test, does anyone even consider the possibility that maybe it's not the test that's flaky, but that the software under test is itself buggy in a non-deterministic way?
I don't have a good answer for how to deal with this. It would be an enormous amount of effort to go through every flaky result and reason about whether the problem is with the test or the code under test, especially in larger codebases. Maybe the heuristic of assuming that it's the test's fault is good enough, and this is what I've done in every team I've worked in.
But it does bother me.
- wyldfire 6y ago> It would be an enormous amount of effort to go through every flaky result and reason about whether the problem is with the test or the code under test, especially in larger codebases. IMO a "flaky test" is one for which the investigation has been performed and the design of the test assessed to be "flaky." i.e. not a test that passes some times and not others, but one whose design cannot reliably expect a given result or one whose implementation can be impacted by factors outside of the control of the test suite. Timeouts are a classic source of flaky tests. Once a test is determined to be flaky it should be evicted from the test suite, or at least demoted from the "reliable/frequently run" rank to a lower one that can tolerate some hand holding.
- nogabebop23 6y agoI'm not sure if it's exhaustive, but any test with a dynamic dependency is one I would likely classify as "flaky", like anything hitting another system. Draw your own boundaries as fit around what "dynamic" means (ex: some might say static files are OK, others not). Also, I'm talking every-commit, unit-level testing for context.
- keegancsmith 6y agoThis is tricky but when you reframe it it's clear the heuristic you mention is the correct one. The value of a test is giving signal to a developer that the code they wrote is correct. If a test is flaky due to the test or the code, the signal the test provides has now become negative. The parent mentions filing it on the team that owns the test. This sounds correct and the onus is then on that team. At larger companies this process is automated. You may enjoy the flake detection section in this blog post https://dropbox.tech/infrastructure/athena-our-automated-build-health-management-system https://dropbox.tech/infrastructure/athena-our-automated-bui...
- hinkley 6y agoThis is why your CI system should record statistics on which tests failed and when. A test with a pattern of failing intermittently over its entire lifetime is most likely a flaky test. Disabling the test modifies a bunch of metrics, like skipped tests and code coverage (woe be unto you if some idiot set it up so you can't commit code that lowers test coverage, instead of just warning you). There are much bigger and more likely sources of regression in the code than a new concurrency bug in code covered by a test that already has a concurrency bug in it. Like someone modifying a test to match the regression they just created. If there is a concurrency or randomness bug in the untested code, that will probably surface when trying to write a robust test to replace the old one. If not by the author, then by the person the author asks for help when they can't figure it out.
- kubanczyk 6y ago> A test with a pattern of failing intermittently over its entire lifetime is most likely a flaky test. I'm not GP, but I think they meant something more radical. In CI, if a test is seen to fail and then succeed on the same git commit for the first time, that's it. One time is enough to classify it as "flaky". No more lifetime, no more patterns of interesting behavior. The test is removed and thrown into backlog.
- hinkley 6y agoI might paint with a broader brush. If your commit breaks master, we revert the whole change and you try again. When the old stuff breaks, there's a bunch of other decisions built on top of it and there is no clear path to unwind the changes. Between HEAD and HEAD-4? Yank it.
- viraptor 6y agoIf you have a problem with testing infrastructure/code, that will start disabling random tests until either you're left with nothing, or realise that radical moves are rarely practical. (With a side of distrust from any team you send the work to)
- dudzik 6y ago> Before disabling the test, does anyone even consider the possibility that maybe it's not the test that's flaky, but that the software under test is itself buggy in a non-deterministic way? I work on the team that build the system in the article. When we disable tests we ask our developers investigate the root cause. Most of the time it is related to application code, so fixing a intermittently failing test definitely isn't related to the test itself, but is often a result of multiple factors.
- smarterclayton 6y agoThere’s one particular domain where this question runs deep - tests that verify parts of distributed systems. Note: please do not extrapolate any problem described below outside of distributed systems. For instance in Kubernetes we run lots of tests that verify parts of the system actually working together, not just mocked. A fair amount of the time, once tests are in and soaked, the test flake is actually the canary that says things like “no, the kernel regressed networking AGAIN” or “no, your retry failure means that etcd lost quorum - which shouldn’t have happened, what on earth is going on, oh, someone suddenly started hotlooping and now the api server is performing 10k writes a second”. At some point, testing the system in isolation is not enough because all failures are interactions of multiple components. Your flakes are the evidence that the gigantic stack of software known as “modern computing run from HEAD lol” is full bugs. It’s definitely a culture change - once anyone anywhere thinks “oh it’s just flaky” they stop treating it like signal. Once they treat it like noise, it’s very hard to unwind. Today, when I look at one of those test suites and see a flake, I assume that everything in open source is broken, again, and pull out the shovel, because 75% of the time it is.
- smarterclayton 6y agoOf course, to be able to see the signal, you have to be ruthless to noise. Deflake first, disable if you must, but never stop until it goes green and then be ruthless about keeping it that way.
- dharmab 6y agoOne of the culture changes I've had to try to push was to test for the failure rate of a distributed system rather than the absolute number of failures- don't test for "Were number of failures less than n", test for "Did the success rate fall below 99.99%"
- haddr 6y agoI totally understand, in our setup, flaky tests happen when testing parts of the software that depend on the Hadoop infrastructure. Sometimes you can’t fully control some pieces and as a result those tests can be flaky. However I want to highlight how frustrating it is when your tests fail when you wait for the build to finish...
- karmakaze 6y agoI have found two cases that occur enough to not assume that the test is wrong: 1. latency/performance - the test system in a well-controlled environment shouldn't have so much latency variance anyway (project was an 'instant' messenger) 2. maybe 30% of the time the test (run frequently in CI) is identifying a race condition that doesn't show up when run ad-hoc by a dev
- greggman3 6y agoI've always been curious about this. My fiction is that if someone spent the time they'd find a few sources of most of the flakiness and and at some point the flakiness would fall drastically. Like maybe there are 1000 flaky tests but fixing 20 bugs would clear up 900 of those flaky tests. No idea if it's true. My experience though is that programmer A who has no domain knowledge of the issue marks test as flaky and files a bug. That bug is completely ignored and there's a giant pile of flaky test and people wondering why stuff broke and it turns out because 1000s of tests are being ignored as flaky
- twic 6y agoWe had exactly this at an old job. A flaky browser test that turned out to be due to some misuse of thread local storage, so occasionally users would see data for another user. Once we tracked that down, we took flaky tests much more seriously.
- nitrogen 6y agoWhen thread-local and request-local aren't the same thing...
- raverbashing 6y agoWell, if it's a third-party software, that's not ideal but "ok" for realistic levels of ok. But yeah, your tests shouldn't be non-deterministic unless there's a very good reason why not (and to be honest I can't think of any) For added resiliency, try changing the order of the tests and/or not running some tests sometimes That being said, my obligatory test rant is "don't write micro-tests" (like "test if method returns" then "test if method returns the right value" etc) and don't set-up everything in the setup function (that is, setup only what you need).
- regularfry 6y agoThe key realisation, I think, is that you've already got that problem. A flaky test is one that's failed to do its job. That's why you can't just disable it, you need to back that up with rework to make sure the underlying problem is addressed.