4 ms·
> This find | xargs mawk | mawk pipeline gets us down to a runtime of about 12 seconds, or about 270MB/sec, which is around 235 times faster than the Hadoop imp
by snaky 8y ago
> This find | xargs mawk | mawk pipeline gets us down to a runtime of about 12 seconds, or about 270MB/sec, which is around 235 times faster than the Hadoop implementation.
https://adamdrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html https://adamdrake.com/command-line-tools-can-be-235x-faster-...
- StreamBright 8y agoWell this is great until you need more nodes. :) I am talking about the same scalability while maintaining a much lower ecological and financial footprint.
- Zariel 8y agoUsing hadoop/spark for <2gb of data seems like a terrible idea. When all you have is a hammer everything starts to look like a nail.