Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rdevulap
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
rdevulap
2mo ago
a simple perf stat should show that the "Keep 50% of random data" case will have insanely more branch mis-predictions that the others.
2.
▲
Copilot implemented a ThreadPool to serve as a replacement for OpenMP
(github.com)
1 points
by
rdevulap
1y ago
|
0 comments
3.
▲
by
rdevulap
2y ago
On similar lines, faster tanh https://github.com/microsoft/onnxruntime/pull/20612
4.
▲
SIMD based custom object and key-value pair sorting in C++
(github.com)
1 points
by
rdevulap
3y ago
|
0 comments
5.
▲
Intel's x86-simd-sort v4.0 Delivers A 2x Boost For AVX-512, Adds AVX2 Code
(phoronix.com)
2 points
by
rdevulap
3y ago
|
1 comments
6.
▲
by
rdevulap
3y ago
Intel's x86-simd-sort 4.0 Delivers A 2x Boost For AVX-512 Performance, Adds AVX2 Code
7.
▲
by
rdevulap
3y ago
Still using the simple way of getting the pivot though.
8.
▲
by
rdevulap
3y ago
Of course! Appreciate all the time you put in. I added a few more optimizations to qsort after that (see https://github.com/intel/x86-simd-sort/pull/33 ), just wanted to know if your analysis took that into ac
9.
▲
by
rdevulap
3y ago
Thank you for such detailed analysis! Just curious to know which version of x86-simd-sort did you benchmark: release v1.0 or the top of main current branch? (I'm the author of x86-simd-sort).