Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dougall
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
dougall
4y ago
Bit late, and the other comments are right, but it's worth noting that pushes typically aren't that expensive. ARM has a 4-byte STP (store paired) instruction that pushes two values at a time. So usually a push only costs two byte
32.
▲
by
dougall
4y ago
Yeah, these features exist, and they help, but I don't think they should be given all the credit. Both "Apple's Secret Extension" and "Total Store Ordering" are features that other emulators can choose to disab
33.
▲
by
dougall
4y ago
Amazing work! It's nice to put a name to it :)
34.
▲
by
dougall
4y ago
Yeah, agreed. I get the impression it's a small team. But there is a long-tail of weird x86 features that are implemented, that give them amazing compatibility, that I regret not mentioning: * 32-bit support for Wine * full x87 emulati
35.
▲
by
dougall
4y ago
Yeah, I haven't either, but I haven't looked. So, I'm not sure, but I wouldn't expect any "function recognition" tricks, since there isn't really static linking, but I would expect the e.g. memcpy and strc
36.
▲
by
dougall
4y ago
In my experience the M1 does have competitive SIMD performance? https://dougallj.wordpress.com/2022/04/01/converting-integer... https://dougallj.wordpress.com/2022/05/22/faster-
37.
▲
by
dougall
4y ago
Yeah, I think Intel's only number so far is "2048 int8 operations/cycle/core" (as opposed to VNNI's 256): https://www.servethehome.com/wp-content/uploads/2021/09/Inte... Whi
38.
▲
by
dougall
4y ago
Yep, still 128-bit.
39.
▲
by
dougall
4y ago
Correct, yep. These are theoretical numbers, measured in cycles from a P-core (with no loads/stores), real-world performance tends to be a little less (~93%): https://twitter.com/stephentyrone/status/145566559
40.
▲
by
dougall
4y ago
This header works with standard Xcode/macOS, by taking advantage of inline assembly in a slightly-cursed way (turning register names into numbers and encoding the instruction itself).
41.
▲
by
dougall
4y ago
Yeah, M2 doesn't. But I think there are some phones using the Snapdragon 8 Gen 1 that have SVE2.
42.
▲
by
dougall
4y ago
On M1, for single-precision, one AMX P-unit is ~1.64 TFLOPs, one P-core is ~102 GFLOPS. So ~16x core-for-core. But you have four P-cores for every AMX P-unit, so more like 4x. And for double-precision that shrinks to 2x (~410 GFLOPs to ~51G
43.
▲
by
dougall
4y ago
I hadn't, but results are here: https://twitter.com/dougallj/status/1561255753339781120
44.
▲
by
dougall
4y ago
Yeah, I haven't looked at the exact changes, but I believe this is the source: https://github.com/apple-oss-distributions/zlib
45.
▲
by
dougall
4y ago
The non-forked zlib hasn't been accepting optimisations: https://github.com/madler/zlib/issues/346 Hopefully changes will be merged from my obscure private fork into four other obscure private forks (zli
46.
▲
by
dougall
4y ago
Nice! (I've been meaning to write up this Apple M1 ~60GB/s version, which I think is similar: https://gist.github.com/dougallj/66151f1c509484a42fe0abd0d84... )
47.
▲
by
dougall
4y ago
Hmm, yeah, this is a bit different, and probably a bit better for Huffman codes (though worse for my x86 use-case, where it's beneficial to know which parts of the subsequent streams are valid before decoding of the preceding stream co
48.
▲
by
dougall
4y ago
Yep. The M1 Max has a single-core memory bandwidth of 102GB/s, and an all-core bandwidth of 224GB/s, so splitting work across cores gives you more bandwidth: https://www.anandtech.com/show/17024/apple-m1
49.
▲
by
dougall
4y ago
The post glosses over it a bit - the CRC32X instruction always uses the common polynomial 0x04C11DB7 (matching zlib, commonly just called CRC-32), and there's a second instruction, CRC32CX, which is the same but uses the polynomial 0x1
50.
▲
by
dougall
4y ago
Hmm, yeah, this might work out... Two SIMD uops process 16 bytes, so each SIMD uop is doing eight bytes of work - the same as CRC32X, but with more frontend pressure (and preferable because they can run on any of the four SIMD ports, not ju
51.
▲
by
dougall
6y ago
I don't really understand the kernel module process/policy - it was a lot of trying random things and seeing what worked. But I think you shouldn't need permission, as long as you're building it and running it on your ow
52.
▲
by
dougall
6y ago
Great post - glad some of that code has been useful! If it's of interest, these performance events (and the whitelist for this API), are described by Apple at https://github.com/apple/darwin-xnu/blob/main
53.
▲
Bitwise conversion of doubles using floating-point multiplication and addition
(dougallj.wordpress.com)
5 points
by
dougall
6y ago
|
0 comments