5 ms·
No one who has done serious work across the database platforms believes they restrict benchmarks because they're behind or uncompetitive. They both make very ca
by defaultname 5y ago
No one who has done serious work across the database platforms believes they restrict benchmarks because they're behind or uncompetitive. They both make very capable, competitive database systems.
The problem, as cited in the linked page, is that "It takes much more work to refute bad benchmarks than to produce them". We see on HN with regularity where people post egregiously flawed benchmarks, usually to demonstrate some preconceived notion or other. And while that can usually be ascribed to ignorance or sloppiness, it's trivial for vendors to contrive benchmarks that are purpose-suited to make their own product look good and the competition look bad, however completely artificial and unrealistic the scenario is.
DeWitt clauses are an abomination. They should not exist. But I get why they exist and are threatened, even if they are basically never actually enforced (seriously though -- are they ever actually enforced?).
- still_grokking 5y agoI think even flawed benchmarks are good benchmarks. Because real world code isn't optimized to the max. Because real world code makes bad or even wrong assumptions. Because there are usually much more requirements to real world code than "be as fast as you can". Additionally, those flawed benchmarks make you think about a lot of interesting details. Details you usually don't think about when producing real world code. Also it's a little bit alike as with statistics. "Don't trust any statistics that you haven't fabricated yourself". Still statistics are useful in general.
- aidenn0 5y agoThere's flawed benchmarks and there are completely wrong benchmarks. I've seen so many benchmarks where the timing or loop overhead dominated to the point that they weren't measuring anything but the cost of taking timing values I've seen benchmarks where they run a single time without warming up the disk cache, so whichever benchmark they happen to run first shows as 2 orders of magnitude slower. I've seen benchmarks where the standard deviation was much larger than the difference between the times reported, but they didn't bother to check for this. I've seen benchmarks where they don't even use the same algorithm, and the difference in the algorithm chosen dominates any other concerns.
- still_grokking 5y agoExactly! And even those very wrong results give quite interesting insides. Such stuff is worthless as benchmarks, sure. But all the other things you can learn form such stuff is sometimes even quite fascinating. So discussing completely wrong benchmarks has imho often some value.
- aidenn0 5y agoIt's entirely possible that the net utility of such benchmarks being published is negative; enough people come away with the initial wrong conclusion without critically investigating it. It's almost certainly true that the net value to the company whose product looks bad in the benchmarks is negative, which is why these clauses exist.
- kaba0 5y agoOf course you can learn from these. But do they describe anything related to the product at hand?
- burnished 5y agoIf the benchmarks were the kind of bad that you allude to in your second statement it would probably be pretty reasonable for the reasons you state, and overall you would probably end up with a distribution of benchmarks that approximated overall performance. What about intentionally crafted, nigh pathological examples? With the kind of flaws that don't crop up in real world code? I don't think those have value. Actually, I think I'm trying to distinguish between flawed benchmarks like you are talking about and actively malignant benchmarks.
- kaba0 5y agoNah, flawed data is pretty much unusable. Like, if I were to benchmark some really short java program by restarting the JVM each time, do I get to say that it is slow? No, because I measured some utterly stupid metric.