Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
clamchowder
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
clamchowder
3y ago
(author here) I think it's not useful to eliminate branches in hot loops. These games have giant instruction footprints and a high branch rate. A loop will probably fit within the L1 BTB and uop cache, and probably won't benefit f
32.
▲
by
clamchowder
3y ago
Maybe it was built from Krait, but I'm not sure. I never tested Krait and don't have microbenchmark code written to deal with 32-bit arm
33.
▲
by
clamchowder
3y ago
(author here) It is not based on A72. Kryo doesn't have anything in common with A72 besides understanding the same instructions. I covered that in detail in the article. Kryo is wider and has a very different out of order engine than A
34.
▲
by
clamchowder
3y ago
(author here) yeah up to 2016. The Snapdragon 821 was released in 2016, and was the last one to have an in-house Qualcomm core
35.
▲
by
clamchowder
3y ago
All modern chips have pipelined decoders, including ARM ones. For example, the Cortex A72 has three decode stages, and it's running a 3-wide decoder at low clock speeds.
36.
▲
by
clamchowder
3y ago
Note that Golden Cove has a 6-wide decoder, and does not have a clear advantage over Zen 4 with a 4-wide decoder. Other parts of the architecture affect performance more, so it doesn't make sense to put more area and power budget into
37.
▲
by
clamchowder
4y ago
(author here) full disclosure, I work at Microsoft. On Azure though, not on Windows. But using Windows, as another commentator pointed out, was done because my latency test gave very weird results from SteamOS and needed more investigation.
38.
▲
by
clamchowder
4y ago
(author of the article here) Agreed. I'm just commenting on the hardware architecture. The deck is a pretty competent gaming device, given its ultraportable form factor and tight power constraints. You can find that out on any number o
39.
▲
by
clamchowder
4y ago
They tell you how to detect support but don't describe the instructions or execution environment (registers, etc).
40.
▲
by
clamchowder
4y ago
Feels like a shotgun approach. They have x86-64 (Zhaoxin), RISC-V (Xiangshan, Alibaba), MIPS -> Loongarch (Loongson), and aarch64 (Phytium). My guess is the central government is trying to build domestic chips by throwing piles of money
41.
▲
by
clamchowder
4y ago
OS is a Chinese linux distro called Loongnix (Loongnix 20, DaoXiangHu). It's Debian based, and we're using binaries from their packages (not compiling from source).
42.
▲
by
clamchowder
4y ago
Yeah you're right about the multiplication performance. I checked back and 64-bit integer multiplication is one per four clocks. I disagree that core count should be taken to mean anything about MT performance. You always have to consi
43.
▲
by
clamchowder
4y ago
(author here) The FPU is not quite equivalent to a Sandy Bridge FPU, but the FPU is one of the strongest parts of the Bulldozer core. Also, iirc multiply throughput is the same on Bulldozer and K10 at 1 per cycle. K10 uses two pipes to hand
44.
▲
by
clamchowder
4y ago
(author here) With hyperthreading/SMT, all of the above are shared, along with OoO execution buffers, integer execution units, and load/store unit. I wouldn't say there are issues with sharing those components. More that the
45.
▲
by
clamchowder
4y ago
Yeah, will have Part 2 up shortly (author here). If you try to make a Wordpress post too long, typing gets extremely laggy to the point of being unusable. So splitting it into two parts lets me go a little more in depth without waiting two
46.
▲
by
clamchowder
4y ago
The problem is loop overhead matters on AMD, because AMD's compiler doesn't unroll the loop. Nvidia's does, so it doesn't matter for them.
47.
▲
by
clamchowder
4y ago
(Author here) See https://github.com/clamchowder/Microbenchmarks/tree/master/G... It's very much a work in progress, as noted in the article. And some of the stuff that worked reasonably well on my
48.
▲
by
clamchowder
4y ago
Apologies for the confusion, the tests were run with ReBAR. I've updated the article to reflect that Shouldn't affect the conclusion for anything besides the PCIe copy to/from GPU tests.
49.
▲
by
clamchowder
4y ago
Yeah, he has some pretty interesting stuff. But AVX512 transition latency is pretty different from clocking up from idle. The CPU is already running at full tilt, and is making a relatively small frequency/voltage change. You can see f
50.
▲
by
clamchowder
4y ago
I meant it's unrelated to how fast a CPU clocks up. If something is taking longer than 0.5 ms, you shouldn't be doing it in the ISR. Queue up a DPC and do your longer running processing there, or send it to user space. And yeah it
51.
▲
by
clamchowder
4y ago
Thanks :) I guess I can't reply to a 7th level comment, so hopefully this one shows up in the right place. I agree, there are multiple factors at play. But I don't think it's basically an implementation choice. Certainly it l
52.
▲
by
clamchowder
4y ago
> I think it's great the author made an account today on HN, and is replying to questions from his post. This is one if the best parts of the community here. :) > I would love to see some test done on newer CPUs So, the site is a
53.
▲
by
clamchowder
4y ago
Yeah, didn't want to start an article with a five paragraph essay especially when wordpress pagination doesn't work, so I can't get an Anandtech style multi-page article up. And yep. You can even run a CPU at full clock all t
54.
▲
by
clamchowder
4y ago
That's a different and unrelated topic. If you're concerned about how fast device driver code can respond, well you can get a lot done in 0.5 ms even with the CPU running at 800 MHz or whatever the idle clock is.
55.
▲
by
clamchowder
4y ago
Big cores each have a 512 KB L2 cache with a latency of around 24 cycles, while the little ones each have a 256 KB L2 with ~23 cycle latency. It's really terrible considering the low 2.34/1.6 GHz clock speeds. Then there's no
56.
▲
by
clamchowder
4y ago
> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually
57.
▲
by
clamchowder
4y ago
Hey, author here. The Snapdragon 670, Snapdragon 821, and i5-6600K tests were done on Linux, and the rest were done on Windows. If Windows is delaying boost, it doesn't seem to be any different from Linux. And lack of intermediate stat