3 ms·
To understand the constrains we have here, a ballpark figure is that we run a million tests per commit across all platforms (this is based on counting test file
by jgraham 10y ago
To understand the constrains we have here, a ballpark figure is that we run a million tests per commit across all platforms (this is based on counting test files, so doesn't account for the fact that some files can contain thousands of separate test functions, but also doesn't account for the fact that we don't run every test job on every push). On average we see around 10 test failures per push due to test instability. So "on average" tests fail about one time in one hundred thousand. Of course, in reality there are some tests that fail much more often than this (one time in a hundred, say) and many that ~never fail.
Whilst test raciness is perhaps the most common problem, we also see intermittent failures due to race conditions in the browser code itself as well as through infrastructure instability. However it's unclear that moving to Rust will make much difference anyhow; Servo still sees intermittent tests so merely eliminating data races in safe code is insufficient to fix this problem.
- chelmertz 10y agoThis is a really interesting problem area. Tests that fails sometimes are really annoying because of the "broken windows" analogy. Are you using the most unstable tests as input of what to redesign next? Is it kind of a "deal with it" situation, where you need to retry the test suite a couple of times per commit, until it becomes green?
- db48x 10y agoThe test suite doesn't get retried, we just wait for another commit to come in; it's not generally a long wait. See https://treeherder.mozilla.org/#/jobs?repo=mozilla-inbound https://treeherder.mozilla.org/#/jobs?repo=mozilla-inbound Another problem is that running the tests takes ages.
- gcp 10y agoYou both "deal with it" and there's task forces that look at the tests that fail most often, and try to find someone with the skills to investigate.
- sethammons 10y agoHow do you rule out that the races are a problem in the code vs a problem in the tests? If we have tests that are sometimes red, we strive hard to remedy them. We've found issues before where we thought it was the tests that were wrong when it was actually the code.
- gcp 10y agoYou can't easily, you have to debug them. Debugging intermittent test failures is hard, and you have no guarantee you're actually improving much tangible things despite the time invested in them. It just sucks.