Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ribit
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
ribit
2y ago
Geekbench supports Intel AMX and AVX-512. This is all in GB documentation.
32.
▲
by
ribit
2y ago
They have not adopted ARMv9. This is still ARMv8, but with SME.
33.
▲
by
ribit
2y ago
Right, so you are disabling all performance features and effectively turning your CPU into a low–end low–power SKU. Of course you’d get better battery life. It’s not the same thing though.
34.
▲
by
ribit
2y ago
M1 Ultra did benchmark close to 3090 in some synthetic gaming tests. The claim was not outlandish, just largely irrelevant for any reasonable purpose. Apple does usually explain their testing methodology and they don’t cheat on benchmarks l
35.
▲
by
ribit
3y ago
You seem to be stuck in the 90-ties. Computing is 64-bit nowadays, not 32-bit, and modern architectures/ABIs integrate frame pointers and frame records in a way that’s both natural and performant.
36.
▲
by
ribit
3y ago
Thanks, this is great! I think this conversation is a great illustration for the idea that data structures can be seen as more fundamental than algorithms.
37.
▲
by
ribit
3y ago
You can implement your B-tree as an array if that's what you like. The advantage of the B-tree-like-layouts is that you need fewer indirections. With binary search you do one node jump per comparison. With a B-tree you do one jump per
38.
▲
by
ribit
3y ago
One also needs to look at the context of these quotes. For example, for the binary search they change the array layout to improve cache locality. This is not really helpful if you have to work with sorted arrays (as many binary search algor
39.
▲
by
ribit
3y ago
Why would you think that replacement part authentication is immoral? Quite in contrary, I’d say it’s an important safety feature for devices that have access to extremely sensitive data. It’s just important that the user has the authorizati
40.
▲
by
ribit
3y ago
I was unable to find any reference to performance or architectural details. Memory bandwidth is very low for an ML-focused device. Price is extremely high. What am I missing?
41.
▲
by
ribit
3y ago
What I had in mind is not the neutral engine but the AMX (Apple's matrix/long vector coprocessor).
42.
▲
by
ribit
3y ago
I am talking about the matrix/vector coprocessor (AMX). You can find some reverse-engineered documentation here: https://github.com/corsix/amx On M3 a singe matrix block can achieve ~ 1TFLOP on DGEMM, I assume it
43.
▲
by
ribit
3y ago
Probably preaching to the choir here, but if you want high-performance sgemm on M-series, you should consider using the system-provided BLAS. It will invoke the dedicated matmul hardware on Apple CPUs, which has access to significantly high
44.
▲
by
ribit
3y ago
Thank you, very insightful and makes perfect sense! I do wonder however why Nvidia and Intel chose not to expose an AXPY/outer product instruction if they use these kinds of operations under the hood. I can imagine them being useful in
45.
▲
by
ribit
3y ago
It is also interesting how this relates to hardware implementations. Nvidia and Intel AMX appear to be using dot product engines under the hood and do a matrix multiplication in a single instruction. Apple AMX and ARM SME use outer product
46.
▲
by
ribit
3y ago
And any of those Mac programs need anything more recent than 4.1? GL is legacy tech and Apple does offer legacy support. I just don’t see any reason to invest more work into this area. We should encourage people to stop using GL, not use it
47.
▲
by
ribit
3y ago
You can implement an OpenGL driver on top of Metal. But why bother dedicating so many resources for the sake of a suboptimal legacy API?
48.
▲
by
ribit
3y ago
Alyssa chooses some very odd language here, it seems to me. Yes, Apple GPUs do not support geometry shaders natively because geometry shaders are a bad design and do not map well to GPU hardware (geometry shaders are known to be slow even o
49.
▲
by
ribit
3y ago
Metal is probably the most streamlined and easiest to use GPU API right now. It's compact, adapts to your needs, and can be intuitively understood by anyone with basic C++ knowledge.
50.
▲
by
ribit
3y ago
It's because Vulkan is designed for driver developers and (to a lesser degree) for middleware engine developers. As far as APIs go, it's pretty much awful. I was very pumped for Vulkan when it was initially announced, but seeing t
51.
▲
by
ribit
3y ago
Thank you, I understand now!
52.
▲
by
ribit
3y ago
I am a bit confused by this since I got agreement with over 90% of people. Are others getting much lower scores? My result suggests that is indeed possible to agree :)
53.
▲
by
ribit
3y ago
Sure, there are things that will always benefit from some ASIC magic (like, I don't see texturing going away any time soon). And we even get new specialized hardware, such as hardware RT. But advances in both performance and rendering
54.
▲
by
ribit
3y ago
Apple has three hardware units for machine learning (if you disregard the regular CPU FP hardware): the neural engine (for energy efficient inference for some types of models), the AMX unit (for low-latency matmul and long vector), and the
55.
▲
by
ribit
3y ago
Hardware rasterizers are likely to disappear before long. We still use them because they are energy and area-efficient when working with large triangles. But already achieving high primitive count is difficult with hardware rasterizers, whi
56.
▲
by
ribit
3y ago
> Don't GPUs also have out of order execution and instruction level parallelism? Not any contemporary mainstream GPU I am aware of. Sure, the way these GPUs are marketed does sound like they have superscalar execution, but if you di
57.
▲
by
ribit
3y ago
It's really great to see these kind of articles! Of course, this is just scratching the surface. I think the most challenging bit about understanding GPUs is breaking through the marketing claims and trying to understand what is really
58.
▲
by
ribit
3y ago
This is quite funny, because I feel the same about what you are saying. The way I understand you is that a seller of the product should grant any third party extensive access to that product, so that the third party can modify and implement
59.
▲
by
ribit
3y ago
> The lack of competition and alternative options is the core issue here. Free market competition will figure it out, that's the simple answer. You know what? I actually agree with this! If alternative stores would really offer comp
60.
▲
by
ribit
3y ago
Right. So you are saying that Meta has the right to implement their services and platforms in any way they see fit, and the user has the right to choose between using those services under Meta's conditions or not using them at all. But
More ›