Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
raphlinus
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
raphlinus
1y ago
You're probably right about this. In the short to medium term, I expect that the Rust and C++ sub-ecosystems will be making different sets of choices. I don't know of any major C++ game or game-adjacent project adopting, say, Dawn
92.
▲
by
raphlinus
1y ago
This is a longer and deeper conversation, but I think on topic for the original article, so I'll go into it a bit. The tl;dr is developer friction. By all means if you're doing a game (or another app with similar build requirement
93.
▲
by
raphlinus
1y ago
Yup, I think slang is the future. Anyone on this thread willing to fund a Rust implementation?
94.
▲
by
raphlinus
1y ago
We tried something like this with piet-gpu-hal. One problem is that spirv-cross is lossy, though gaps in target language support are getting better. For example, a device scoped barrier is just dropped on the floor before Metal 3.2. Atomics
95.
▲
by
raphlinus
1y ago
So I would say skill at GPU assembly is in-demand for the elite tier of GPU performance work. Not necessarily writing much of it (though see [1] for an example, this is the kernel of multisplit as used in Nvidia's Onesweep implementati
96.
▲
by
raphlinus
1y ago
The question of which assembly is best to learn is of course incredibly subjective, but I think the author gives short shrift to ARM32. It is historically important (especially for the Acorn computers, most popular in the UK), sensibly desi
97.
▲
by
raphlinus
1y ago
I read LLVM (or one of its many GPU-flavored variants) reasonably often, mostly to figure out where in the chain a shader miscompilation is happening. But I've never personally had to write it, and it's not easy for me to think of
98.
▲
by
raphlinus
1y ago
Thanks so much, Peter, for writing this up. I think it adds a lot to the record about what exactly happened with the Cell. And, as with Larrabee, I have to wonder, what would an alternative universe look like if Sony had executed well? Or i
99.
▲
by
raphlinus
2y ago
Thanks for posting this, I'll take a look. It wasn't on my radar, but the idea of doing a DSL specifically for SIMD is something I've been thinking about and also starting to explore myself.
100.
▲
Towards fearless SIMD, 7 years later
(linebender.org)
177 points
by
raphlinus
2y ago
|
175 comments
101.
▲
by
raphlinus
2y ago
There are a lot of odd exclusions on that list. Just spot-checking, I see blog.plover.com, the blog of Mark Jason Dominus, who by the way is looking for a job[1]. Also, dtrace.org is excluded, which hosts four individual blogs that surely s
102.
▲
by
raphlinus
2y ago
Yup. A little more detail on the overheating part in particular is here: https://github.com/AsahiLinux/speakersafetyd
103.
▲
by
raphlinus
2y ago
Getting reasonable speaker support in Asahi Linux was a big deal. Part of the problem is that limiting the power usage to prevent overheating requires sophisticated DSP. Without that, you get very limited volume output within safe limits. P
104.
▲
by
raphlinus
2y ago
Regarding the SIMD optimizations, the authors may want to look into faer. I haven't had a great experience with its underlying library pulp, as I'm trying to things that go beyond its linear algebra roots, but if the goal is prima
105.
▲
by
raphlinus
2y ago
Absolutely. And the fact that we need to evolve both is one of the reasons progress has been difficult.
106.
▲
by
raphlinus
2y ago
I consider Xeon Phi to be the shipping version of Larrabee. I've updated the post to mention it.
107.
▲
by
raphlinus
2y ago
> It is an instance of Larrabee in the same sense as AMD Zen 4 is an instance of Larrabee. This is an odd claim. Clearly Xeon Phi is the shipping version of Larrabee, while Zen 4 is a completely different chip design that happens to run
108.
▲
by
raphlinus
2y ago
The problems I'm having are very different than those for raytracing. Sure, it's dynamic, but at a fine granularity, so the problems you run into are divergence, and often also wanting function pointers, which don't work well
109.
▲
by
raphlinus
2y ago
Possibly compilation and linking. That's very slow for big programs like Chromium. There's really interesting work on GPU compilers (co-dfns and Voetter's work). Optimization problems like scheduling and circuit routing. Sear
110.
▲
I want a good parallel computer
(raphlinus.github.io)
233 points
by
raphlinus
2y ago
|
194 comments
111.
▲
by
raphlinus
2y ago
The argument I have in mind is subtle and nuanced, and I didn't write clearly in that comment (the bit about the smart solo programmer was mostly sarcasm but with a grain of truth). But to try to answer: The value of UB is to clearly
112.
▲
by
raphlinus
2y ago
Oh hey, I also have "in defense of undefined behavior" in the queue of blog posts I'd like to write some time, with that exact title. What a coincidence. That said, it's unlikely to get written as I have things that are
113.
▲
by
raphlinus
2y ago
Well, I have some qualifications in typography, a reasonable familiarity with ML techniques, and am fairly good at math (though I can't claim to have won the Putnam), and I can imagine how to apply ML for this task.
114.
▲
by
raphlinus
2y ago
Agree with sibling comments. There's something very slippery and tricky going on with "perceptual area," it's not simple geometry. This is actually an area where I think machine learning has something to offer.
115.
▲
by
raphlinus
2y ago
Yup, nothing wrong with clear exposition about simpler algorithms, there's definitely a place for that. I just thought HN readers should have some more context on whether we were looking at a programming exercise or state of the art al
116.
▲
by
raphlinus
2y ago
The second one is Thomas Smith's independent reimplementation of Onesweep. For the official version, see https://github.com/NVIDIA/cccl . The Onesweep implementation is in cub/cub/agent/agent_radix_
117.
▲
by
raphlinus
2y ago
This is not a fast way to sort on GPU. The fastest known sorting algorithm on CUDA is Onesweep, which uses a lot of sophisticated techniques to take advantage of GPU-style parallelism and work around its limitations. Linebender is working (
118.
▲
by
raphlinus
2y ago
You might enjoy this talk by Erik Lindholm (now retired), who talks about Riva 128 and many of the other early Nvidia cards: https://ubc.ca.panopto.com/Panopto/Pages/Viewer.aspx?id=880a...
119.
▲
Closing the “green gap”: energy savings from the math of the landscape function
(terrytao.wordpress.com)
111 points
by
raphlinus
2y ago
|
71 comments
120.
▲
by
raphlinus
2y ago
I just want to say I'm rooting for you, and hope you enjoy the book and learn a lot from it. I had a bad experience with complex analysis as a teen (took a grad class that was a bit over my head). Many years later, I got Tristan Needha
More ›