4 ms·
> The tools used for benchmarking were BenchmarkTools.jl for Julia, microbenchmark for R, and timeit for Python. Is this a meaningful comparison? These benchma
by nnm 6y ago
> The tools used for benchmarking were BenchmarkTools.jl for Julia, microbenchmark for R, and timeit for Python.
Is this a meaningful comparison? These benchmarks repeatedly call the function (with same input file) many times and then take the average of the time spent. However, in real world, we only read a specific file ONCE -- if we call the csv reader several times, it is almost always for different files.
- oizin 6y agoI think this is an important point, if you look at the warmup run on [1] the Julia CSV readers are slower than many of the R and python packages . It makes the claims of 10-20x faster quite disengenuous in my opinion. A lot of data science work will only read the file on once per session and what these benchmarks actually suggest for that sort of work is not reflected in the title. [1] https://www.queryverse.org/benchmarks/ https://www.queryverse.org/benchmarks/
- StefanKarpinski 6y agoThere are two potential issues that this might disregard: 1. cold file cache 2. JIT compile time The cold/hot file cache issue affects all languages equally, so it doesn’t invalidate the comparison. It is possible that CSV file is not in cache and a system's disk I/O is slow enough that it becomes a bottleneck, which would make all parsers equally slow. However, this is not very realistic because CSV files—especially large enough ones where you care about speed—are almost always compressed—So you're not going to be bottlenecked on disk read. Assuming that compressed disk read + decompression is fast enough to keep the data flow high, you're back to CSV parsing being the bottleneck. The JIT compile cache only affects Julia. The reason is it not included in the benchmark results is because it is a small, fixed overhead: it is only paid on the first CSV file read (no matter the size), and if you read a larger data set, the compile time does not increase. Since the point of benchmarks is typically to project from a smaller case how long it would take to do even larger tasks, you don't want to include small, fixed overheads. For some concrete numbers, I just timed reading a tiny CSV file on my MacBook Pro 2018 and the first read, including compile time took 4 seconds. The second read took 0.000347 seconds. So that's the fixed overhead we're talking about here for reading a CSV file: about 4 seconds on the very first CSV file you read. People can be the judge of whether that's a showstopper for them or not.
- ChrisRackauckas 6y ago4 seconds if one doesn't use PackageCompiler, but I'd argue using PackageCompiler for something like this (and Plots) is fairly standard now. I probably update my basic package compiles once every month (which is often because I work on a lot of packages of course), and many of the SciML users report updating PackageCompiler sysimages every few months. With that, the basic compile times are gone. Given that as a "new standard", we might as well include some times with compile times in a user-improved environment.
- nnm 6y agoThanks for the explanation: now it seems to me the comparison is not perfect but still meaningful.