4 ms·
The test suite is ultra important. Even in the unlikely case that you have an engineer who thoroughly understands the machine and language specs, you have a ton
by neeeeees 6y ago
The test suite is ultra important. Even in the unlikely case that you have an engineer who thoroughly understands the machine and language specs, you have a ton of existing, useful, programs that rely on under-defined behaviour in one or both of them.
- mehrdadn 6y agoAh, I might have to tease out my question here. By correctness I'm only referring to the implementation being true to the language specification, not to "not breaking existing programs". So if a program relies on undefined behavior and then gets broken in the next version of the compiler because of that behavior, I don't consider that a correctness problem in the compiler for the purposes of this question. (Which is not to suggest keeping it working isn't a worthwhile goal. I'm saying that's not the aspect of it that I'm asking about.)
- pjmlp 6y agoThe problem is that many don't learn C and C++ properly, they have read some language introduction book and assume "how my compiler does it" is the right way and probably never touched a copy of ISO. So plenty of programs rely on UB even without realizing it.
- fit2rule 6y agoMore to the point, anyone doing compiler validation work would a) always have a full test suite as their first line of attack, b) always have a well understood developer pipeline in place, to ensure a consistent understanding - coding rules, code review processes, etc. and c) always looks at what the compiler is actually doing, compared to what they think its doing - i.e. is able to navigate code -> assembly and all the stages in between, with competence. Anything less is asking for trouble.
- joosters 6y agoA ‘full’ test suite? I’m not sure that one of those has ever existed for any large program, let alone one under active development. We have to work in the real world here!
- fit2rule 6y agoYes, a full test suite. 100% code coverage is a thing, even in compiler validation. Also, every test that was ever written for a compiler is still out there, being run, somewhere. And, for the record, the real world is pretty big.
- andrewaylett 6y agoBut 100% code coverage isn't necessarily a useful thing in this context, as different code (and different options) mean different interactions between passes. And you're never going to cover all the different paths through the compiler. Back when I worked on a commercial compiler, we had a smoke test which we ran before committing. It would take ~30 minutes to run. We had a longer, roughly eight hour, set of tests which we ran daily. And we had machines running tests 24/7, which managed to clock up roughly 20 million distinct test cases (combinations of code and options) over the course of a year. Every single compiler bug we'd ever seen was present (in a reduced form) in the conformance tests we'd run before release. That took about a week.
- fit2rule 6y ago>And you're never going to cover all the different paths through the compiler. Depends who you work for and how willing they are to let such justifications stand in a court of law.
- joosters 6y ago100% code coverage does not mean a full test suite. You can still find huge numbers of bugs in code that has 100% coverage in a test suite. The two are not the same!
- 6y ago
- alkonaut 6y agoI have absolutely no problem finding out that we have a dozen new bugs in production for our 15 year old codebase, because we upgraded our compiler and it revealed places we used undefined or unspecified behavior. I’d much rather have that than a compiler that evolves slower because its developers worry about existing non-conforming code. I can always use the old version of the compiler if I don’t want it to break.
- klodolph 6y agoThe compiler doesn’t do a good job of revealing undefined behavior, it usually just makes your program break in mysterious ways.
- alkonaut 6y agoA better example is perhaps behavior that shouldn’t be relied upon: a standard library unordered collection (Set) where the implementation for 10 years preserves insertion order but says in the docs not to rely on it. Lots of programs have deliberately or not relied on it. Then suddenly it changes impl to one that doesn’t preserve insertion order for slightly better perf. You don’t want the std library developers to worry about existing code as long as the documentation always said “don’t rely on insertion order”. Same with a list sort that is stable but never guaranteed it would be, then turns unstable much later. I was bitten by this one in .NET as were many others, and as a response the developers added hacks into the code to simulate the old behavior it detected the code was built against the older lib. That didn’t make the problem easier to find...
- saagarjha 6y ago> You don’t want the std library developers to worry about existing code as long as the documentation always said “don’t rely on insertion order”. Sometime practicality wins out and the change is reverted when it breaks everything. It’s nice to come up with purist views, and in this case being able to point to documentation and say that the entire world needs to change is the morally high stance to take, but it’s not always what happens. (Some may say that this is what needs to happen to undefined behavior, but that one’s still up in the air. As it so happens I am on the “read the standard” camp…)
- perlgeek 6y agoThen the test suite is still hugely important -- you just have to design it to reflect what you consider "true to the language specification". Or phrased differently, the test suite is (or should be) the machine-readable language specification.
- barrkel 6y agoThat is only theoretical correctness, not useful correctness.