3 ms·
> We’re talking about 31 million lines of data here (all lines with Swiss email addresses in the leak). On our workstation, reducing this to the distinct email
by aw3c2 8y ago
> We’re talking about 31 million lines of data here (all lines with Swiss email addresses in the leak). On our workstation, reducing this to the distinct email addresses (approx. 3,3 million) took a mere 10.099 seconds (4 + 6 + .099).
That's just 700MB of data... On my 5 year old middle-class PC this took 12 seconds with sort and uniq.
$ wc -l email30m
31000000 email30m
$ time sort email30m | uniq > email30m.uniq
real 0m11.498s
user 0m33.833s
sys 0m0.933s
$ wc -l email30m.uniq
2172320 email30m.uniq
This automatically split the task into all available CPU cores.
- dahfizz 8y agoUsing Unix utilities is so underrated. I love Unix programming for the simplicity and power that you demonstrate well here. It's not as flashy and cool as writing some big program but I am constantly amazed how people can go through new web framework fads every 2 months but these 40-50 year old programs are still the most useful tools around.
- munro 8y agoI'm getting 2 minutes and 29.98 seconds, with my 2017 MacBook Pro with 8 cores + 16 GiB ram. I'm using GNU tools, could you share more details on how you're getting 12 seconds? $ time awk 'BEGIN{for(i=0; i<3100000; i+=1) for(y=0; y<10; y+= 1) print i "@gmail.com"}' > email30m.sorted awk > email30m.sorted 8.29s user 0.82s system 97% cpu 9.352 total $ ls -lh email30m.sorted 522M Mar 8 17:14 email30m I didn't have your email30m dataset, so I fabricated one, mine is smaller at 522M and has less entropy, since 'm assuming everyone uses gmail & uses numbers for usernames. :) $ time shuf email30m.sorted > email30m shuf email30m.sorted > email30m 15.58s user 1.69s system 98% cpu 17.561 total I also shuffled the dataset, because that would be cheating. $ wc -l email30m 31000000 email30m $ time sort email30m | uniq > email30m.uniq sort email30m 873.55s user 4.29s system 585% cpu 2:29.98 total uniq > email30m.uniq 4.10s user 0.34s system 2% cpu 2:29.98 total $ wc -l email30m.uniq 3100000 email30m.uniq I followed what you did, but it took way longer when running on my machine. $ time sort email30m.sorted | uniq > email30m.sorted.uniq sort email30m.sorted 413.75s user 3.66s system 468% cpu 1:29.04 total uniq > email30m.sorted.uniq 4.11s user 0.40s system 5% cpu 1:29.04 total And just out of curiosity, it takes 1 minute 29 seconds to sort & uniq the presorted dataset.
- vthriller 8y agoGNU sort generates a bunch of temporary files for large inputs, and for a lot of linux folks /tmp is mounted as tmpfs (i.e. it's RAM/swap-backed), but it might do something else on other platforms, or just locate temporary files on disk or something, so that might be one explanation. Or it could be the difference between qsort() from different libcs.
- mappu 8y agoI ran all your commands above on Debian Buster (4-core i5 2500 from 2011, 8GB ram, SATA SSD). My /tmp apparently isn't a tmpfs mount. Results: http://paste.debian.net/1072430/ http://paste.debian.net/1072430/ It took 54s to sort|uniq and 38s for the presorted dataset. EDIT: For the unsorted dataset: 35s for `sort --parallel=8 -u` and 37s for `sort --parallel=4 -u`.
- aw3c2 8y agohttp://0x0.st/zHb2.7z http://0x0.st/zHb2.7z All I did was place the file in a ramdisk.
- vthriller 8y ago> This automatically split the task into all available CPU cores. GNU sort can already utilize multiple cores, and you can override internal heuristics about whether it's going to be beneficial for one particular input or not with --parallel=N. Anyways, if multithreaded sort is available what you should probably do is replace `sort | uniq` with just `sort -u` to save some time on pipe IO.
- aw3c2 8y agoOh nice! `sort -u` is just as fast as `sort | uniq` for me but good call :)
- stewbrew 8y agoTo be fair, you describe the last step of one of his analyses. He starts with 900gb of data. One could of course discuss whether spark is the right tool for filtering a text file for a pattern, IMHO the post is interesting nevertheless.
- antpls 8y agoIt's nice to see that Spark is on par with Unix tools on low data volumes. The Spark program can be applied as-is for data volume 10 times or 100 times bigger without being modified. You would have to reinvent Spark if you started from Unix tools and applied it to bigger volume of data