4 ms·
And the 5 years old AMD Ryzen 5 5600H is doing 7M? Am I reading this right? Then I need to try this on Strix Halo
by throwa356262 2mo ago
And the 5 years old AMD Ryzen 5 5600H is doing 7M?
Am I reading this right? Then I need to try this on Strix Halo
- rbanffy 2mo agoAnd it's only using AVX-2 and not AVX-512, AMX or ACE. Or built-in GPUs and NPUs (the M series doesn't emphasize matrix multiplication on the CPU side because it already has matrix multiplication units on the GPU, which is always attached).
- bigyabai 2mo agoBefore the M5, there was no dedicated matrix multiplication hardware on the Apple Silicon GPU. Their solution was generally using the NPU and AMX coprocessors for tensor and matrix workloads.
- saidnooneever 2mo ago[dead]
- ranger_danger 2mo agoBut do processors actually offload any CPU opcodes to their GPU? That could be quite useful if it can be used to improve execution speed.
- rbanffy 2mo agoNo, but if you are guaranteed to have a GPU or NPU packaged together it becomes like the SPUs in the Cell processor - you can’t offload instructions (unless you use a trap mechanism to a subroutine) but you can have code that hides the setup, the different ISA, and the result retrieval, behind an API call.
- throwa356262 2mo agoThe NPU in previous AMD generation has unfortunately it's own dedicated RAM that is far too small for an LLM. But even those should comfortably run an SLM of around 100-400M parameters at crazy high speeds and with minimal power usage.
- rbanffy 2mo ago> an SLM of around 100-400M parameters at crazy high speeds and with minimal power usage Software support for that kind of usage, when available, will be immensely impactful.