3 ms·
Sometimes there's a middle ground: make your "map" and "reduce" steps separate scripts. If you want to do the parsing in Python instead of awk, just make a tin
by ims 8y ago
Sometimes there's a middle ground: make your "map" and "reduce" steps separate scripts.
If you want to do the parsing in Python instead of awk, just make a tiny script that reads from stdin and writes to stdout - that way you can put it between xargs or parallel and whatever else is in the pipeline.
The parallelization is a separate concern, so it doesn't need to be mixed in with the parsing (or whatever) concern. The downloading is a separate concern; use wget or requests in a Python script or whatever, it doesn't need to be mingled with the parsing.