4 ms·
Tinygrad targets consumer hardware (to be precise, only Radeon 7900XTX and nothing else[1]), while ROCm does not actually provide good support for such hardware
by Lockal 2y ago
Tinygrad targets consumer hardware (to be precise, only Radeon 7900XTX and nothing else[1]), while ROCm does not actually provide good support for such hardware. For example, last release of hipBLASLt-6.1.1 library has deep integration with PyTorch[1], while working only on AMD Instinct hardware. And even for the professional hardware out there, the support period is ridiculous: AMD Instinct MI100 (2020) is not supported. Only 4 years and tens of thousands of dollars worth of hardware is going to the trash, yay!
And to be more precise, they still use some core libraries from ROCm stack[3], they just don't use all these fancy multi-gigabyte[4] hardware-limited rocBLAS/hipBLASlt/rocWMMA/rocRAND/etc. libraries.
[1] https://tinygrad.org/#tinybox https://tinygrad.org/#tinybox
[2] https://github.com/pytorch/pytorch/issues/119081 https://github.com/pytorch/pytorch/issues/119081
[3] https://github.com/tinygrad/tinygrad/blob/v0.9.0/tinygrad/runtime/autogen/hsa.py#L388 https://github.com/tinygrad/tinygrad/blob/v0.9.0/tinygrad/ru...
[4] https://repo.radeon.com/rocm/yum/6.1.1/main/ https://repo.radeon.com/rocm/yum/6.1.1/main/
- anthonix1 2y agoEven without hipBLASlt, PyTorch is still ~4x faster than tinygrad on a 7900 XTX for GPT2, and works fine. Any idea why?
- Lockal 2y agoMaybe due to 2x slower matmul (well known bottleneck), as one of multiple factors? https://github.com/tinygrad/tinygrad/blob/v0.9.0/extra/gemm/hip_matmul.py https://github.com/tinygrad/tinygrad/blob/v0.9.0/extra/gemm/... There are multiple bounties just for it in https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic1qUWzpfAzuD_J06N1JwOCGs/edit#gid=0 https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic...
- anthonix1 2y agoI think the matmul issue is symptomatic of a much deeper issue. It would be nice to see less whining and blaming AMD (PyTorch and llm.c actually work on 7900 XTX, and blow tiny grad out of the water in terms of perf!), and more just getting stuff to work.
- wozeparrot 2y agoThe idea with the experimental backends is that they will talk directly with the kernel drivers, [3] is for the HSA backend but is not needed for the AMD backend.