Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
janwas
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
janwas
1y ago
hm. Doesn't the existence of Vulkan subgroups and CUDA shuffle/ballot poke huge holes in their 'SIMT' model? From where I sit, that looks a lot like SIMD. The only difference seems to be that SIMT professes to hide (or u
62.
▲
by
janwas
1y ago
On vqsort: yes, the current RVV set of shuffles is awfully limited and several implementations produce one element per cycle. We also saw excessive VSETVLI, though I understand that has been fixed by an extra compiler pass. Could be interes
63.
▲
by
janwas
1y ago
I advised the Abseil design and regret not pointing this out earlier: changing the interface to insert/query batches of items would be considerably more efficient, especially for databases. Whenever possible, 'vertical' algor
64.
▲
by
janwas
1y ago
(For other readers:) This is what our Highway library does - wrapper functions around intrinsics, plus a (constexpr if possible) Lanes() function to query the length. For very many cases, writing the code once for an 'unknown to the pr
65.
▲
by
janwas
2y ago
Or how about https://www.notebookcheck.net/Way-to-run-DeepSeek-s-671B-AI-... 768 GiB for $6000.
66.
▲
by
janwas
2y ago
Or also getauxval? Highway has code to check for this, including that vectors are at least 128 bits: https://github.com/google/highway/blob/master/hwy/targets.cc...
67.
▲
by
janwas
2y ago
Amen! I really do not understand this. It has been 7 years since SVE was introduced. Writing an application in terms of a specific lane count loses performance portability - either it's too many, or too few, for the particular CPU. And
68.
▲
by
janwas
2y ago
Yep, including draining the store buffer. We've gotten our ThreadPool barrier+wait to use only acq/rel, but not yet the work stealing. Does anyone have experience with that already?
69.
▲
by
janwas
2y ago
Yes indeed, this is what we do :) There is an opaque Mask type for which operations such as CountTrue, AllFalse etc are provided. If you really want the one or other representation, VecFromMask and BitsFromMask/StoreMaskBits convert as
70.
▲
by
janwas
2y ago
> except for byte-level processing, variable-length codecs, or mixed-precision numerics. That never works with autovectorization and can’t be solved with general-purpose SIMD wrappers. Counterexamples: Chromium's byte-level HTML sca
71.
▲
by
janwas
2y ago
The auto-vectorization (which I anyway would not rely on) default setting also sounds like a workaround for the SKX issue. For atomic, I'm curious how you make use of that?
72.
▲
by
janwas
2y ago
In addition to ISPC, it is possible to do this kind of vector-length abstraction at the library level, e.g. in our Highway library. We routinely write code that works on 128-512 bit vectors. Some use cases are harder than others, e.g. trans
73.
▲
by
janwas
2y ago
Strongly agree with the first part of your post :) BTW in addition to the weights, it's also interesting to consider the precision of accumulation. f16 is just not enough for the large matrix sizes we are now seeing. (Gemma.cpp TL here
74.
▲
by
janwas
2y ago
Perhaps the use cases are different (heavily data-parallel), but FWIW I do not remember many cases where we were frontend bound, so icache hasn't been a concern.
75.
▲
by
janwas
2y ago
hm, fair enough. IIRC JPEG XL was a few hundred KB of SIMD code for the four or so different targets/ISAs, including the generic fallback, but I can believe video codecs are larger.
76.
▲
by
janwas
2y ago
How can tuning be independent of devising the algorithm? Are you really suggesting writing a variant of a kernel, tuning it to the max, then discovering a new and different way to do it, and then discarding the first implementation? That se
77.
▲
by
janwas
2y ago
Given the wider availability of masking (AVX-512, RISC-V and SVE), I figure scalar tails are no longer the preferred pattern everywhere.
78.
▲
by
janwas
2y ago
Intrinsics have the huge advantage of enabling wrapper functions, which remove the ugly names and allow you to write user code only once, such that it is even portable (or at least multiplatform-dependent). Good point about asan and other i
79.
▲
by
janwas
2y ago
Collaborators have actually superoptimized some of the more complicated Highway ops on RISC-V, with interesting gains, but I think the approach would struggle with largish tasks/algorithms?
80.
▲
by
janwas
2y ago
Var-width SIMD can mostly be written using the exact same Highway code, we just have to be careful to avoid things like arrays of vectors and sizeof(vector). It can be more complicated to write things which are vector-length dependent, such
81.
▲
by
janwas
2y ago
I'm also in the mission-critical camp, with perhaps an interesting counterpoint. If we're focusing on small details (or drowning in incidental complexity), it can be harder to see algorithmic optimizations. Or the friction of chan
82.
▲
by
janwas
2y ago
I'm curious why there are even function calls in time-critical code, shouldn't just about everything be inlined there? And if it's not time-critical, why are we interested in the savings from a custom calling convention?
83.
▲
by
janwas
2y ago
hm, I'd be concerned about relying on autovectorization. How much better is 'better'? Compiler friends have told me that something permute-heavy like sorting is unlikely to soon work, if ever. My biased opinion, from doing th
84.
▲
by
janwas
2y ago
Interesting that both de-novo and porting seems to have worked. I do not understand why GGML is written this way, though. So much duplication, one variant per instruction set. Our Gemma.cpp only requires a single backend written using Highw
85.
▲
by
janwas
2y ago
> Python is only a thin veneer over those backends. There is an interesting writeup by dvyukov that shows the potential cost of "thin veneers" :) https://github.com/dvyukov/perf-load
86.
▲
by
janwas
2y ago
Valid points, but I am doing exactly that full-time for the past years :) Including compression and quicksort, which are nontrivial uses of SIMD. The 2x128 is handled by adopting AVX2 semantics for our InterleaveLower/Upper op, and whe
87.
▲
by
janwas
2y ago
Lemire and collaborators often write in C++ intrinsics, or thin platform-specific wrappers on top of them. I count ~8 different implementations [1], which demonstrates considerable commitment :) Personally, I prefer to write once with porta
88.
▲
by
janwas
2y ago
Note that the article mentions using both outputs of the instruction, whereas the emulation is only able to compute one output efficiently.
89.
▲
by
janwas
2y ago
I agree unrolling is usually helpful :)
90.
▲
by
janwas
2y ago
We have a different philosophy: not supporting/encouraging needlessly SIMD-hostile software. We assume users properly allocate their data, for example using the allocator we provide. It is easy to deal with 2K aliasing in the allocator
More ›