4 ms·
> This is an unreasonable standard when you do not know in advance how big the output is. Why is that unreasonable? Let us say a single job outputs 10% of the
by MainJane 5y ago
> This is an unreasonable standard when you do not know in advance how big the output is.
Why is that unreasonable?
Let us say a single job outputs 10% of the free space. As long as you run fewer than 10 jobs in parallel, GNU paralel can run forever, because it spits out the output when a job is done and then frees up the space for this job, while starting the next one.
A simple example:
yes 1000000 | parallel -j10 seq | pv >/dev/null
On my laptop I get 600 MB/s which would fill /tmp in a few minutes, and it does not.
When dealing with big data it is not uncommon that the total data piped between commands is way larger than the free space on /tmp (which is typically fast, where as free space on $HOME is slow - thus setting $TMPDIR to $HOME/tmp may slow down your job drastically).
If you only have 5 minutes, I hope you will use them on providing actual code to support your claim, that "The comparison is not very fair to modern day xargs."
If it takes longer than 5 minutes to code, I would say your use of "easy-ish" is unwarranted.
You leave me with the feeling that you have not thought this through and that the reason why you do not provide any code is because you are now realizing you are wrong, but you do not have the guts to admit so.
Prove me wrong by posting the code. It should be "easy-ish" :)
You can use this as the test case to implement:
yes 1000000 | parallel -kj10 "echo 'This is double spaced '{#}; seq {}" | pv >/dev/null
- cb321 4y agoYou are just moving goalposts from "grouping to not mix" in the comparison doc to "grouping to not mix with exact space management profile(s) of GNU parallel". Even worse, you now bring in IO space-speed assumptions, other use cases (hay generation not needle search), various dissembling and childish "taunts for proof" when you clearly understood the suggestion well enough to analyze it for potential limitations. Your attitude is the problem, not missing code. Also, I never said "/tmp" and the paths could be FIFOs with record size/buffering limitations instead. Speaking of /tmp filling and questionable space management defaults: yes 2000000000 | parallel seq | pv > /dev/null fills my /tmp disk partition (or $TMPDIR) before emitting one byte to pv with invisible (unlinked) temp files. Not ideal. GNU sort at least shows me there are files present yet also seems to clean up on Ctrl-C. There is likely some solution to fix this in 15 kLOC of gross Perl. I did not find it in "5 minutes" (another unreasonable standard since the many 1000s of lines of GNU parallel docs take far longer to read, but you already seem to ignore my explanations of "unreasonable"). You even anticipate this in your 10% example. At least in my life, "way more" is often much more than 10x more. So, you basically contradict yourself. As to the actual subtopic, besides being unfair/out-of-date, the comparison tableau is also incomplete - maybe willfully so, as per too common marketing dishonesty. "Proof?" People use parallelism to speed things up and need to make decisions about job granularity to not have perf killed by overhead. Some would say this matters more than 95% of the tableau evaluation points. Yet, no overhead benchmarks. Maybe they make GNU parallel look bad?
- MainJane 4y agoIf you feel I am "moving the goalposts" why not just prove your original case? If you are spending 5 minutes on reading the source code, why not instead spend them on proving your original assertion is correct? You can then let the readers decide if they feel I "move the goalposts". I included the example: yes 1000000 | parallel -kj10 "echo 'This is double spaced '{#}; seq {}" | pv >/dev/null to give you some fixed "goalposts" to aim for: Provide a solution that gives the same output byte for byte. Also you do not seem to get the point about the amount of data. I regularly have output from a single job that is bigger than RAM, but rarely have output from a single job that would fill /tmp. However, the total combined output from all the jobs will often take up more space than /tmp. In numbers: RAM=32 GB, /tmp=400 GB, a single job=33 GB, number of jobs=1000, jobs in parallel=8. In other words: Running all jobs and saving the outputs into files before outputting data will not be useful for me. If you want to use FIFOs I really cannot see how you can deal with output that is bigger than RAM, unless you mix output from different jobs - which again would not be useful to me. But prove me wrong by spending 5 minutes on building the solution. As for your example: yes 2000000000 | parallel seq | pv > /dev/null How would you design this, if output from different jobs are not allowed to mix? If they are allowed to mix paralel gives you: # bytes are allowed to mix yes 2000000000 | parallel -u seq | pv > /dev/null # only full lines are allowed to mix yes 2000000000 | parallel --lb seq | pv > /dev/null none of these use space in /tmp. I sit back with the feeling you are willing to spend hours complaining, but not 5 minutes on proving your assertion that it can be done "easy-ish". Prove me wrong: Spend 5 minutes on the task you believed was "easy-ish". If it cannot be done in 5 minutes, be brave enough to admit you were wrong.