8 ms·
How Rust is tested
- JoshTriplett 9y ago> Today the longest-running configuration takes over 2 hours. I'm curious which configuration that is, and how it could be accelerated.
- sanxiyn 9y agoFYI, it's i686-apple-darwin at the moment.
- steveklabnik 9y agoLooking at the build that's building right now, all of them have passed but one, which has been going on for 2 hours and 23 minutes: https://travis-ci.org/rust-lang/rust/builds/252052639 https://travis-ci.org/rust-lang/rust/builds/252052639 > "env": "RUST_CHECK_TARGET=dist RUST_CONFIGURE_ARGS=\"--target=aarch64-apple-ios,armv7-apple-ios,armv7s-apple-ios,i386-apple-ios,x86_64-apple-ios --enable-extended --enable-sanitizers --enable-profiler\" SRC=. DEPLOY=1 RUSTC_RETRY_LINKER_ON_SEGFAULT=1 SCCACHE_ERROR_LOG=/tmp/sccache.log MACOSX_DEPLOYMENT_TARGET=10.7\n", iOS, it seems. Looking at the other builds, * s390x Linux * i686-apple-darwin * x86_64-apple-darwin Seems apple is an overall slowpoke.
- sanxiyn 9y agoIt seems to vary from build to build... Compare https://travis-ci.org/rust-lang/rust/builds/251809226 https://travis-ci.org/rust-lang/rust/builds/251809226 where everything is under 2 hours except i686-apple-darwin.
- ihsw2 9y agoThis has more to do with Travis-CI than Darwin et al. https://www.traviscistatus.com/ https://www.traviscistatus.com/ Notice how the active OS X builds reaches the limit of 216 jobs and the backlog of OS X jobs increases without decreasing. This is a capacity issue that is likely related to high OS X build costs. Another thing to notice is that the backlog correlates strongly with time of day in the EST timezone, where the backlog starts to take off after 9am EST.
- ihsw2 9y agoAs mentioned elsewhere in this thread, it stands to reason that Travis-CI is the culprit here. https://www.traviscistatus.com/ https://www.traviscistatus.com/ Their slowest build jobs are all on Darwin/OSX-based systems and (notably) this is a problem numerous other projects have -- slow jobs on OSX build systems. There is a large backlog of OSX build jobs that starts around 9am EST.
- JoshTriplett 9y agoIs that true for the paid version of Travis as well? If not, might be worth paying for it, to support a community as large as Rust's. (If the paid version is that backlogged as well, I'm surprised they don't put more resources into accelerating it; perhaps not enough of their customers care about macOS?) Alternatively, given that Mozilla has their own extensive CI infrastructure (for Firefox), perhaps it'd make sense to migrate to that instead of Travis?
- steveklabnik 9y agoWe are already paying Travis. > Given that Mozilla has their own extensive CI infrastructure While Mozilla does sponsor a lot of Rust stuff, it's fundamentally a community project, and so many of us (including myself) would prefer to not to move closer to something that's Mozilla-specific.
- JoshTriplett 9y agoThanks for the explanations. > We are already paying Travis. Have you heard anything from them about improving the issues with macOS? You've certainly got a case study in it being a problem. > While Mozilla does sponsor a lot of Rust stuff, it's fundamentally a community project, and so many of us (including myself) would prefer to not to move closer to something that's Mozilla-specific. Fair enough; I only suggested it because I had the impression that Mozilla's bits were all based on Open Source systems, and it would just be a matter of sharing hosting infrastructure rather than duplicating it.
- 9y ago
- JoshTriplett 9y ago> At some point though LLVM began doing a valid optimization that valgrind was unable to recognize as valid, and that made its application useless for us, or at least too difficult to maintain. What optimization was that? Some searching didn't turn up any relevant results (other than this article).
- steveklabnik 9y agoI couldn't quite remember either, and brson told me that https://github.com/rust-lang/rust/issues/11710 https://github.com/rust-lang/rust/issues/11710 was probably it. Can't seem to find where it was actually finally disabled though.
- kibwen 9y agoMoney quotes: "This is a known valgrind false positive and Rust already has a lot of suppressions for it. It's reported upstream on the LLVM bug tracker but is not really an LLVM issue, just a limitation of valgrind. LLVM is allowed to generate undefined reads if it can not change the program behavior." [...] "Just for the record, the undefined read doesn't change the behaviour because it's just a check whether free should not be called, i.e. either it jumps over the free call right away, or it performs the second check."
- pornel 9y agoI appreciate the thoroughness of this. I've been using Rust for 2 years now, and the compiler has been rock solid.
- losvedir 9y agoI love the concept of the "crater run" (or I guess now "cargo bomb"?), where the entire crate (i.e. "package") ecosystem is compiled, as a test for speculative compiler changes. It's one of the huge benefits of being both a statically-typed compiled language and offering a modern package system. A lot of languages have the types & compiling but no standard package ecosystem (Haskell, Go, Java), or the package ecosystem but not the compiling (Ruby+Gems, JS+npm, etc). Rust is the only language I can think of that has both, although I'm sure there are others. I'm not sure if there's tooling for compiler writers in other languages to compile them all, though. And this feature will only get more and more powerful as the ecosystem grows. I think it's a seriously important tool to keep the rust language feeling stable.
- steveklabnik 9y ago> "crater run" (or I guess now "cargo bomb"?) Crater is the older tool that runs builds. CargoBomb is the newer tool that runs builds and tests.
- mrec 9y agoYeah, I was really impressed by that as well when it was first announced. Makes perfect sense once you hear it, but as an old fogey I just wasn't used to thinking on that kind of scale. I had a similar lightbulb moment years ago when someone pointed out that y'know, these days it's perfectly feasible to test a simple float32 function by running every possible float32 bitpattern through it.
- amichal 9y agoTrue... as long as you know what output you are expecting for all 2^32 patterns (Not your point I know)
- Dylan16807 9y agoThese days? Grab a 20 year old Pentium II and you can test a 300-cycle 32-bit function in an hour.
- houli 9y ago
- Scaevolus 9y ago> There’s a big downside though in that landing patches to Rust is serialized on running the test suite on every patch, and it takes a particularly long time to Run the Rust test suite in all the configurations we care about. One solution is speculative batched testing. Kubernetes uses this to keep up with its very high rate of PRs (>40 merges/day) and ~1hr maximum test times. Given a queue of patches A, B, C, D, start the normal testing against A, but also start testing the merge of A+B+C+D against the master. If the batch passes first, merge them all. If A is merged first, it's still "compatible" with A+B+C+D, so they can go ahead. Doing a batch in parallel with a single PR like this slightly increases testing load, but ideally doesn't cause any slowdowns versus the fully serialized case. It looks like bors already supports something like this with "rollups", where simple fixes can be marked for testing together: https://internals.rust-lang.org/t/batched-merge-rollup-feature-has-landed-on-bors/1019 https://internals.rust-lang.org/t/batched-merge-rollup-featu...
- brson 9y agoOh, neat. I've never heard it proposed quite that way before, where both branches being tested are compatible. That's real smart. So far we haven't been able to stomach doing parallel integration builds because of the doubling of expenses. As you see, our current solution is rollups, where a human batches low-risk PRs into one big PR. It works just ok, and is super high-maintenance.
- Scaevolus 9y agoYou already trigger Travis tests for PRs before bors does an auto run, right? In that case, it wouldn't double expenses. Assuming batches are slightly effective, it should reduce expenses-- with >50% success rate with batches of >=3 PRs, you don't have to do individual testing for the later PRs in the batch.
- brson 9y agoThe Travis smoke tests against PRs are only done in a single configuration, so are relatively cheap. I think it would double expenses because our CI runs at capacity. To do parallel builds we would have to contract for double the compute resources. (With a different purchase structure for our CPU time you may be right about that, but not sure).
- hsivonen 9y agoIt would be really nice if Travis had Linux-on-ARMv7 and Linux-on-aarch64 options on real ARM hardware. I wonder what the economics of Travis building such a thing on top of an existing ARM cloud offering like Scaleway would be.
- nercury 9y agoI don't get it, because you have just explained how to unit test everything else. Which is way better than nothing. Not sure how it is moot, also not sure how it is related to the article :)
- jorgec 9y agoUnit test in a nutshell, for a business system: Practically every business system consists of a view, the logic and the database, plus some webservice and whatnot, but mainly its those 3 parts (the so called 3 layers). Can you unit test, test the view layer? Not really. Can you unit test, test the database by doing real test? Again, not really and its not always possible. So, Unit test is mainly focused in the layer in between of the view and database and both, by their nature, can't be automatically tested. I like the idea of Unit Test but, for business system, its moot.
- wfunction 9y ago> The thing we do differently from most is that we run the full test suite against every patch, as if it were merged to master, before committing it into the master branch, whereas most CI setups test after committing, or if they do run tests against every PR, they do so before merging, leaving open the possibility of regressions introduced during the merge. Hm, Travis CI on GitHub runs tests on a pull request before the merge. Is what they're doing really that unusual?
- saghm 9y agoAs it says in the quote you provided, the difference is that it tests what the result would be of the merge, which differs in that a merge might introduce regressions; if you've never had a bug introduced by a merge not doing what you expected, consider yourself one of the lucky ones.
- wfunction 9y agoI'm not sure I'm understanding things correctly. What I meant was that the test chronologically occurs before the merge takes place. But the test includes the merge itself; it's testing what the result would be after the merge. Have you used the system I'm talking about/do you know what I mean? Am I still misunderstanding something?
- tyoverby 9y agoHere's an example scenario: 1. Your PR is tested with the merge against Master. All green! 2. Another PR is tested with the merge against Master. All green! 3. The other PR goes in first. All good! 4. Your PR goes in (merge happens, but introduces a bug). Oh no!
- wfunction 9y agoThanks! Someone pointed it out in a sibling comment afterwards too. Kind of sucks, I wish they had a "do final check and merge" button that would merge PRs sequentially into master after re-testing when there's a collision like this. Any idea why they don't?
- sudeepj 9y agoI have been taking inspiration from the Rust project and the way they go about things in general. In my workplace, we have a bot similar to bors, which does the merge request testing and is the gatekeeper. Tools like bors act as a "force-multiplier".
- crncosta 9y agoI dream with the day we will have a GCC Rust compiler.
- pas 9y agoWhat would the benefits of that be?
- TotallyGod 9y agoHeads up that your Google fonts (in site.css) trigger a warning about insecure scripts, and in Chrome at least they therefore don't load by default.
- Outrageous 9y agoAnyone care to enlighten me why fuzz testing and not mutation testing?