2 ms·
I did some testing on the same (kind of) dataset and task: First test: A single 2.9GB file time rg Result all.pgn | sort --radixsort | uniq -c 13 [Result "
by maxmunzel 7y ago
I did some testing on the same (kind of) dataset and task:
First test: A single 2.9GB file
time rg Result all.pgn | sort --radixsort | uniq -c
13 [Result ""]
1106547 [Result "0-1"]
1377248 [Result "1-0"]
1077663 [Result "1/2-1/2"]
rg Result all.pgn 1.12s user 0.55s system 99% cpu 1.680 total
sort --radixsort 3.87s user 0.37s system 71% cpu 5.911 total
uniq -c 2.69s user 0.02s system 45% cpu 5.909 total
Using Apache Flink and a naive implementation It took 13.969 seconds.
Second test: same dataset, split between 4 files
time rg Result chessdata/ | awk -F ':' '{print $2}' - | sort --radixsort | uniq -c
13 [Result ""]
1106547 [Result "0-1"]
1377248 [Result "1-0"]
1077663 [Result "1/2-1/2"]
rg Result chessdata/ 1.70s user 0.97s system 42% cpu 6.292 total
awk -F ':' '{print $2}' - 5.47s user 0.07s system 88% cpu 6.289 total
sort --radixsort 4.13s user 0.42s system 43% cpu 10.559 total
uniq -c 2.73s user 0.03s system 26% cpu 10.559 total
Flink: 12.724s
Conclusion: For this kind of workload, both approaches have comparable runtimes, even tough taco bell programming has the upper hand (as is should for simply filtering a text file). It took me about equally long to implement both. I think both approaches have their use case.
I ran this locally on my Laptop with 4 logical cores.
- ma2rten 7y agoHadoop is very slow, because it persist the data to disk before every stage. You really wouldn't want to use Hadoop if you don't have a good reason too. More modern tools like Spark and Flink fare better there.