4 ms·
Most of the well-know "unix-style" command-line tools, such as grep, sort, etc. actually have very high performance. Their relatively constrained use cases allo
by robbles 13y ago
Most of the well-know "unix-style" command-line tools, such as grep, sort, etc. actually have very high performance. Their relatively constrained use cases allow the authors to implement decent algorithms and optimizations (e.g. sort uses merge sort, grep uses all kinds of optimizations: http://lists.freebsd.org/pipermail/freebsd-current/2010-August/019310.html http://lists.freebsd.org/pipermail/freebsd-current/2010-Augu...)
In contrast, when you're building a custom pipeline in a high-level language, you're optimizing for simple solutions and are not likely to get better performance unless you hit an edge case where the standard tools do really poorly.
- paulrosenzweig 13y agoActually sort is even better than normal (purely in-memory) merge sort. It looks at available memory and writes out sorted files to merge. http://vkundeti.blogspot.com/2008/03/tech-algorithmic-details-of-unix-sort.html http://vkundeti.blogspot.com/2008/03/tech-algorithmic-detail...
- robbles 13y agoInteresting! So if I understand that article correctly, it's basically doing a multi-phase merge sort where each individual run is stored in a file?