3 ms·
Hi. I made the graph. It was definitely not my intent to obscure or distort any information. Here's the raw data if you like [1]. Some context: the input text i
by felixhandte 8y ago
Hi. I made the graph. It was definitely not my intent to obscure or distort any information. Here's the raw data if you like [1]. Some context: the input text is silesia.tar, which is a standard mixed corpus of data. It was benchmarked by @Cyan4973 on his i7-9700k.
I guess I'm confused about how you feel we're misrepresenting the data? Are you talking about the sentence preceding that graph, "The benefits we’ve found typically range from a 30 percent better ratio to 3x better speed"? That's more describing the benefits we've seen in the real world, where for a variety of reasons (many of which are discussed in the rest of the post), we generally can get more out of zstd than this vanilla benchmark shows.
[1] https://gist.github.com/felixhandte/f6a91bf775d6e7df76ce06c0a8866f91 https://gist.github.com/felixhandte/f6a91bf775d6e7df76ce06c0...
- hinkley 8y agoEdward Tufte will set you straight. The first rule of objective graphs is always show the origin at 0. As an informed graph consumer, if the origin is not 0 you should immediately distrust your eyes, and question the motives of the person showing you the data. The relative sizes are being distorted. Why? In this case, there's a 15% difference in Y values in your data that on the graph is presented as two lines that are separated from each other by 33%. Just by removing 0 and 1 from the chart. On the other hand, that effect is diluted a bit by what you've done on the X axis. Log scale tells a story about trends. The relative slope of two lines on log scale says something. The distance between them doesn't mean much, if anything, although the brain can't help trying to make it mean something. I believe you will find that a great deal of what you lose by plotting y = 0 you'll gain back by plotting x on a linear scale. Those lines deserve to be much farther apart.
- felixhandte 8y agoThe Visual Display of Quantitative Information is one of my favorite books! 0 in this case is not really a relevant value (since that would mean transforming the input into something infinitely large). The functional identity value / origin here is 1x. Here's what that looks like [1]. To me, this is a significantly less useful image. But maybe that stems from a lot of comfort with both the subject matter and the detail log-log plots that I work with to evaluate zstd performance, e.g. [2]. [1] https://imgur.com/gU2Gdf6 https://imgur.com/gU2Gdf6 [2] https://github.com/facebook/zstd/pull/1317#issuecomment-426032139 https://github.com/facebook/zstd/pull/1317#issuecomment-4260...
- hinkley 8y agoHey thanks for the new chart, and you're right about 1x. But we might have to disagree because I don't think this graph is that bad. I can see a bunch of talking points in this graph that are about strategy instead of about why I hate bad graphs. First, that data point for lz4 out around 820 MB/s kind of throws the groove off of that graph. Despite that, I would still suggest you prune the graph off at 850 or 900 to reduce the squish of the horizontal data (pruning the graph to put the last value at 100% of width or height isn't considered a no-no). One of the things I can see now but couldn't before is that the size/speed tradeoff is pretty linear except for the giant dog-leg around 3.6:1, and a less pronounced but still notable one occurs at 2.9:1. If I sat down with my team to discuss this chart I'd suggest we agree that we aren't interested in anything above 3.6:1. And then I'd suggest we look at everything above 2.9:1, but my eye is on 3.2:1 (where zlib tops out and is 1/10th the speed). If they don't like 2.9:1, then the next interesting point is at about 2.7:1 when zlib bottoms out. We're already at such a high bandwidth rate that something else is probably going to be the bottleneck.