3 ms·
Even without hipBLASlt, PyTorch is still ~4x faster than tinygrad on a 7900 XTX for GPT2, and works fine. Any idea why?
by anthonix1 2y ago
Even without hipBLASlt, PyTorch is still ~4x faster than tinygrad on a 7900 XTX for GPT2, and works fine. Any idea why?
- Lockal 2y agoMaybe due to 2x slower matmul (well known bottleneck), as one of multiple factors? https://github.com/tinygrad/tinygrad/blob/v0.9.0/extra/gemm/hip_matmul.py https://github.com/tinygrad/tinygrad/blob/v0.9.0/extra/gemm/... There are multiple bounties just for it in https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic1qUWzpfAzuD_J06N1JwOCGs/edit#gid=0 https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic...
- anthonix1 2y agoI think the matmul issue is symptomatic of a much deeper issue. It would be nice to see less whining and blaming AMD (PyTorch and llm.c actually work on 7900 XTX, and blow tiny grad out of the water in terms of perf!), and more just getting stuff to work.