3 ms·
When all you have is a hammer, every problem starts looking like a nail. The basic premise is fine: If you have a simple problem, using simple tools will give
by reagent_finder 7y ago
When all you have is a hammer, every problem starts looking like a nail.
The basic premise is fine: If you have a simple problem, using simple tools will give you a good result. Here you have text files, you just want to iterate through them and find a result from ONE line that's the same in every file, collate the results. No further analysis required.
Every problem in the world can be solved by a bash one-liner, right!?
There's an interesting dichotomy with bash scripts: One school says any bash script over 100 lines should be rewritten in Python, because it's overcomplex already. Another school says any Python script used daily over 100 lines should be rewritten in bash so there are no delusions about it being easy to maintain.
The original article is from 2013, and doesn't try to do any optimization (I guess, the original article is unavailable at the time of writing of this comment), so it would be an interesting question to see what you could do at the Hadoop end to make the query faster. I would imagine quite a lot.