3 ms·
I'm curious whether you've got any interesting projects underway to work on the flaky test issue? I've been looking at FBs test lifecycle approach with some in
by IanWhalen 11y ago
I'm curious whether you've got any interesting projects underway to work on the flaky test issue? I've been looking at FBs test lifecycle approach with some interest (http://goo.gl/QtUNxY http://goo.gl/QtUNxY), but would love to hear about any other approaches to solving the same problem.
And FWIW on building MongoDB with Evergreen (https://evergreen.mongodb.com https://evergreen.mongodb.com) we've found that a stepback approach seems to give us the finest granularity with minimal cost in identifying the introduction of an error - we batch commits for execution, but when a specific test fails we start running just that test on each previous commit until we hit a passing execution. It obviously doesn't work perfectly in the face of flaky tests (see above question) but it seems to do pretty well.
- phasmantistes 11y agoWe are working on a few things. The first is automatic retries of fast tests. If a test runs quickly and fails, it costs us little to try again just to make sure. Most of our unit tests are configured to run up to 3 times. Another thing is keeping track of a database of individual test case passes/failures across all time. This will let us automatically mark tests as flaky if they fail often, and ignore their results programatically rather than requiring a human to manually mark the test as ignorable. A third thing is, obviously, automatically filing bugs against the owners/authors of tests which have been marked flaky. This is controversial -- often a test is just fine until one of its underlying libraries has a race condition introduced, and the real person to fix it should be the author of that change, not the author of the test. But it is still a step in the right direction much of the time. Many people subscribe to the philosophy that "a flaky test is worse than no test", because you think it is giving you information when in fact it is giving you none. I subscribe to a slightly different philosophy: "A test with a known flaky rate is hugely valuable". If you know how often a test flakes (statistically), then you can measure variances from that rate to detect changes. Of course, a flaky test with an unknown rate of flaky is still useless. Hence the second initiative above: measuring the rate of flake of everything.
- cpeterso 11y agoMozilla has a big problem with flaky tests (aka "intermittents") too. The SpiderMonkey VM can run tests in a deterministic mode to make tests less flaky by making things like GC more predictable. Firefox has an experimental "chaos mode" that takes the opposite approach. It purposely randomizes behavior by adjusting thread priorities, changing hash table iteration order, and randomize timer durations. Unfortunately, many flaky tests fail in chaos mode, so it is not enabled by default. http://robert.ocallahan.org/2014/03/introducing-chaos-mode.html http://robert.ocallahan.org/2014/03/introducing-chaos-mode.h...
- jgraham 11y agoBut we (Mozilla) are also doing many of the same things as the Chromium team here. In particular there is work in progress to automatically "ignore" the results of known-flaky tests until we detect that there has been a change in the rate of flakiness, at which point we will — assuming all goes to plan — trigger new test runs until we can determine the point at which the regression was introduced. I think one of the lessons we've learnt is that with a browser-type project it's very hard to make test runs fully deterministic, for both technical and human reasons. The technical reasons are touched on in the original article: these are complex codebases with lots of moving parts and lots of environmental dependencies. Of course there are various tactics to try and combat this; for example there is a wiki page dedicated to innocuous-looking code that leads to intermittent tests [1]. The human reasons centre around the difficulty of getting people to care about spending time fixing a test that fails one time in 1,000 (which is still very noticeable when you are running it hundreds of times a day). Unless the issue is something that fits a known pattern it's hard work, difficult to tell if your fix even worked, and not likely to be considered a top priority due to the diffuse, hard to quantify, nature of the benefits. I think the fact that both Google and Mozilla still have significant problems with intermittents despite talented engineering staff and it having been a known problem for years implies that some of the standard thinking about making tests fully deterministic simply doesn't apply; for this kind of work you have to embrace — or at least accept — the randomness, and look for ways to get the data you need despite the noise. [1] https://developer.mozilla.org/en-US/docs/Mozilla/QA/Avoiding_intermittent_oranges https://developer.mozilla.org/en-US/docs/Mozilla/QA/Avoiding...
- bgirard 11y agoAt Mozilla we're actually currently working to run some of our tests continuously under RR (http://rr-project.org/ http://rr-project.org/), our record-replay debugger. When a flaky test fails the replay can carefully studied to understand why test only fails say 1 out of 10,000 runs.