Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
janwas
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
janwas
5mo ago
?? Where did you see mention of AI?
32.
▲
by
janwas
5mo ago
Thanks for sharing. The first link seems non public indeed. I can imagine there is some compile issue we could reasonably fix, with the help of someone who has Z13 access. Please encourage them to raise an issue. I will be back on May 26. A
33.
▲
by
janwas
5mo ago
Fair point. If it helps, our security team has called Highway critical infrastructure and helped to harden the repo. The flip side of standardization is that it would be much harder and slower to add ops as the need arises, which we do regu
34.
▲
by
janwas
5mo ago
:) I figure there is always something left to improve. For some kernels which really want to keep 30+ live registers, the compiler might not do as good a job as careful manual tuning, so intrinsics can have a bit of a cost. But I also figur
35.
▲
by
janwas
5mo ago
Yes, the EMU128 target is scalar only, with for loops. This is a fun way to see how well autovectorization works, with the same source code. That works on any CPU. Curious which projects have such concerns, any link?
36.
▲
by
janwas
5mo ago
In such discussions, whenever you mention abstractions are universally "pretty poor", to the extent anyone is listening, I think this hyperbole can do real damage. Maybe it prevents people from getting relevant performance gains
37.
▲
by
janwas
5mo ago
This works today :) Highway provides such an abstraction for arbitrary vector lengths and maps them to intrinsics. All on the library level, no need to wait years for compiler or language updates.
38.
▲
by
janwas
5mo ago
:) I agree a tutorial would be helpful. We are working on one with Fastcode.
39.
▲
by
janwas
5mo ago
Have you considered our Highway library? Runtime dispatch need not be a PITA :) It's basically portable intrinsics, and a much more complete set (>300) than the ~50 in std.
40.
▲
by
janwas
7mo ago
Highway TL here. I agree with the main points, with a few clarifications: > tag-dispatched free functions like hn::Mul(d, a, b) We only require tags for certain ops, mainly memory, casts and reduction; not arithmetic. Operator overloadin
41.
▲
by
janwas
7mo ago
Looks like the ratification plan for Zvzip is November. So maybe 3y until HW is actually usable? That's a neat trick with wmacc, congrats. But still, half the speed for quite a fundamental operation that has been heavily used in other
42.
▲
by
janwas
7mo ago
(Personal opinion) I get the impression that RISC-V-related discussions often lack of awareness of prior work/alternatives. A large amount of (x86) software actually uses our Highway library to run on whatever size vectors and instruc
43.
▲
by
janwas
9mo ago
:D Your code was nicely written and it was a pleasure to port to SIMD because it was already very data-parallel.
44.
▲
by
janwas
9mo ago
Gemma.cpp has nested thread pools, one per chiplet, and one across all chiplets. With such core counts it is quite important to minimize any kind of sharing, even RMW atomics.
45.
▲
by
janwas
11mo ago
> performance of general solutions without using SIMD, is good enough too, since all of which will eventually dump right down to the uops anyway, with deep pipeline, branch predictor, superscalar and speculative execution doing their mag
46.
▲
by
janwas
1y ago
Wow, that number requires STRONG caveats, lest it be called out as completely false. Take away the tensor cores (unless you only do matmuls?), and an H100 has roughly 2x as many f32 flops as a Zen5 CPU, which is considerably cheaper. I susp
47.
▲
by
janwas
1y ago
CPU-time would over-emphasize regions where many threads are running, right? I find wall-time useful for finding serial regions that aren't yet parallelized. More detail here: https://github.com/dvyukov/perf-load .
48.
▲
by
janwas
1y ago
:) Yes indeed, it's about 500 LOC in https://github.com/google/highway/blob/master/hwy/ops/generi... .
49.
▲
by
janwas
1y ago
I appreciate your efforts to nudge readers towards SoA data structures and varying SIMD widths. FWIW I have observed that advice is more effective if communicated with some kindness.
50.
▲
by
janwas
1y ago
Thanks for expanding on your viewpoint. > Why would writing an optimizing compiler qualify as territory for directly writing SIMD code, but anything else is off the table? I understood "directly writing" to mean assembly or eve
51.
▲
by
janwas
1y ago
You also get the automatic support for newer instructions (and multiversioning) with a wrapper library such as our Highway :)
52.
▲
by
janwas
1y ago
Highway author here :) I'm curious what you disagree with, because it all sounds very sensible to me?
53.
▲
by
janwas
1y ago
Makes sense :) Generic or fallback versions are also useful for correctness testing and benchmarking.
54.
▲
by
janwas
1y ago
I made the same argument a while ago but a coworker changed my mind. Can you afford to write and maintain a codepath per ISA (knowing that more keep coming, including RVV, LASX and HVX), to squeeze out the last X%? Is there no higher-impa
55.
▲
by
janwas
1y ago
Mature such wrappers exist, for example our Highway library :)
56.
▲
by
janwas
1y ago
If SIMT is so obviously the right path, why have just about all GPU vendors and standards reinvented SIMD, calling it subgroups (Vulkan), __shfl_sync (CUDA), work group/sub-group (OpenCL), wave intrinsics (HLSL), I think also simdgroup
57.
▲
by
janwas
1y ago
I don't understand why it helps to "avoid them" entirely. For the (in my experience) >90% of shared code, we can gain the convenience of the wrapper library. For the rest, Highway allows target-specific specializations ami
58.
▲
by
janwas
1y ago
I know dzaima is aware, but for all the other posters who might not be, our Highway library provides all these missing instructions, via emulation if required. I do not understand why folks are still making do with direct use of intrinsics
59.
▲
by
janwas
1y ago
Nice work :) Clang x86 indeed unrolls, which is good. But setting the CC and AA mask constants looks fairly expensive compared to fixed-pattern shuffles. Yes, the 2D aspect of the sorting network complicates things. Transposing is already h
60.
▲
by
janwas
1y ago
:) Yes, vqsort recently tickled a bug in clang. I've seen a steady stream of issues, many caused by SLP or the seeming absence of CI. You might try re-enabling it on GCC. Yes, the issue with the sorting network is that it is limited to
More ›