Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jkool702
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
jkool702
6mo ago
> I'm curious about the performance is of forkrun "echo ." in a billion jobs vs. say pure C Short answer: in its fastest mode, forkrun gets very close to the practical dispatch limit for this kind of workload. A tight C lo
2.
▲
by
jkool702
6mo ago
So I thought about this for a bit, and this actually doesnt surprise me all that much. This makes sense when you consider the following 2 things: First, 14k items in batches of 100 are only 140 batches. 140 batches in 160 ms is not even 100
3.
▲
by
jkool702
6mo ago
I appreciate the high praise re: forkrun. forkrun's NUMA approach is really largely based on the idea that, as you said, "real workloads mostly wait for memory". The waiting for memory gets worse in NUMA because accessing mem
4.
▲
by
jkool702
6mo ago
Im happy to hear forkrun is working well for you! I'll have to look into what would be required to package forkrun for the various distros. I'll try to make it happen in the near-ish future.
5.
▲
by
jkool702
6mo ago
I hate to say it but forkrun probably wont work in cygwin. I haven't tried it, but forkrun makes heavy use of linux-only syscalls trhat I suspect arent available in cygwin. forkrun might work under WSL2, as its my understanding WSL
6.
▲
by
jkool702
6mo ago
forkrun complements things like SLURM (and even MPI). forkrun is intra-node, and is all about utilizing all the resources any given node as efficiently as possible, including when the node has a deep NUMA topology (e.g., it's EPYC-base
7.
▲
by
jkool702
6mo ago
whoopsie... code_to_parallelize "$nn" should be code_to_parallelize "$nn" &
8.
▲
by
jkool702
6mo ago
The difference you are seeing in this specific usage is because frun is dynamically adjusting batch size and worker count (by default it always begins at a batch size of 1 and using 1 worker). It is pretty darn good at dynamically pinning t
9.
▲
by
jkool702
6mo ago
parallel works fine so long as the time per job is on the order of seconds or longer. Let me give you an example of a "worst-case" scenario for parallel. Start by making a file on a tmpfs with 10 million newlines yes $'
10.
▲
by
jkool702
6mo ago
> And if you’re objective, what could be done to other tools to make them competitive? I wanted to reply separately to this bit, because I needed a bit of time to think about and respond to it. To be frank, parallel optimizes for "b
11.
▲
by
jkool702
6mo ago
10 years represents going from maxJobs=$(nprocs) while read -r nn; do code_to_parallelize "$nn" (( $(jobs -p | wc -l) > maxJobs )) && wait -n done < inputs to a NUMA-Aware Contention-Free
12.
▲
by
jkool702
6mo ago
> I ask because I see multiple em dashes in your description here, and a lot of no X, no Y... notation that Codex seems to be fond of. I asked a few LLM's for tips on writing the HN post. The post is my own words, but their style ma
13.
▲
by
jkool702
6mo ago
curl isnt required - you just need to source the `frun.bash` file. Downloading frun.bash and sourcing it works just fine. directly sourcing a curl stream that grabs frun.bash from the github repo is just an alternate approach. It is not &qu
14.
▲
by
jkool702
6mo ago
How did it work for you?
15.
▲
by
jkool702
6mo ago
Theres no "install" - you just need to source the `frun.bash` file. Downloading frun.bash and sourcing it works just fine. directly sourcing a curl stream that grabs frun.bash from the git repo is just an alternate approach. It is
16.
▲
by
jkool702
6mo ago
So, in forkruns development there have been a few "AHA!" moments. Most of them were accompanied by a full re-write (current forkrun is v3). The 1st AHA, and the basis for the original forkrun, was that you could eliminate a HUGE a
17.
▲
by
jkool702
6mo ago
So...yes, the execve overhead is real. BUT there's still a lot you can accomplish with pure bash builtins (which don't have the execve overhead). And, if you're open to rewriting things (which would probably be required to so
18.
▲
by
jkool702
6mo ago
So, there are a few reasons why forkrun might work better than this, depending on the situation: 1. if what you want to run is built to be called from a shell (including multi-step shell functions) and not Go. This is the main appeal of f
19.
▲
by
jkool702
6mo ago
Hi HN, Have you ever run GNU Parallel on a powerful machine just to find one core pegged at 100% while the rest sit mostly idle? I hit that wall...so I built forkrun. forkrun is a self-tuning, drop-in replacement for GNU Parallel (and xargs
20.
▲
Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
(github.com)
151 points
by
jkool702
6mo ago
|
41 comments
21.
▲
by
jkool702
1y ago
no problem. thanks for helping me discover and fix a bug in timep that all my test cases missed. Sorry the overhead is too high (relative to cube.bash's insanely low avg command runtime of something like 1 microsecond) to be really use
22.
▲
by
jkool702
1y ago
> Thanks. I suppose this will depend on each script, as there is another commenter here claiming that the overhead is much higher. The better way to think about overhead with timep is "average overhead per command run" (or more
23.
▲
by
jkool702
1y ago
It took some time, but I figured out what was causing the errors when profiling cube.bash - the code assigns huge (some >400,000 elements) associative arrays in a single command, and timep was taking the full (several MB) $BASH_COMMAND f
24.
▲
by
jkool702
1y ago
also re: BATS Im aware of it, but have never ended up actually using it. Ive heard before the sentiment you imply - that its great for fairly simple script...but not so much for long and complicated scripts. And, well, the handful of bash p
25.
▲
by
jkool702
1y ago
the LINENO is (mostly) reliable, so long as the code that is running comes from a file somewhere. timep runs the code that you want profiled by generating a script file that: 1. declares all the variables timep uses to track state and thing
26.
▲
by
jkool702
1y ago
Yay, a comment! >I find it hard to believe that there is minimal overhead from the instrumentation, as the README claims. Surely the additional traps alone introduce an overhead, and tracking the elapsed time of each operation even more
27.
▲
Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
(github.com)
53 points
by
jkool702
1y ago
|
12 comments
28.
▲
by
jkool702
1y ago
Im currently working on adding the ability to record user/sys cpu time (in addition to wall-clock time) to timep. This will be coupled with a new modification to the timep_flamegraph.pl script that will control the flamegraph coloring
29.
▲
Show HN: Timep – a next-gen profiler and flamegraph-generator for bash code
(github.com)
26 points
by
jkool702
1y ago
|
1 comments
30.
▲
Show HN: Timep – a next-gen time-profiler and flamegraph-generator for bash code
(github.com)
3 points
by
jkool702
1y ago
|
0 comments
More ›