Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anthonix1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
anthonix1
2y ago
... which also has a much lower power cap
2.
▲
by
anthonix1
2y ago
Yeah I would suggest taking a look at PyTorch on AMD before saying stuff like "scaled_dot_product_attention is an NVIDIA CUDA kernel exposed as a PyTorch function", because that is demonstrably false. Also, FWIW, I would suggest g
3.
▲
by
anthonix1
2y ago
Any direct comparisons to 8xH100? 2 toks/sec seems very slow! I haven't done any LoRA training on MI300x myself, but I have done LLama 3.1 full training on 8xMI300x and got pretty close to 8xH100 performance with my own kernels
4.
▲
by
anthonix1
2y ago
Does JAX have its own implementations of matmul, flash attention etc? Or does it use the ROCm implementations like PyTorch does? (e.g,. hipblaslt, Composable Kernel FA etc) Not too familiar with JAX, but the abysmal PyTorch training perf on
5.
▲
by
anthonix1
2y ago
Do they support curvilinear cells?
6.
▲
by
anthonix1
2y ago
AHh gotcha. Well yeah I reckon you render a full custom 4004 w/ koru patterned transistors into about 4m x 4m stained glass panel. Would look good as the foyer panel for the CS dept at the University of Waikato
7.
▲
by
anthonix1
2y ago
Don't bother with the rectilinear pakeha layouts, do your half adders in curvilinear patterns, Koru style
8.
▲
by
anthonix1
2y ago
OK, so in the case of llm.c, if you're just including the HIP headers, using hipblasLt, etc, what would be the benefit of using scale instead of hipify?
9.
▲
by
anthonix1
2y ago
I ported Karparthy's llm.c repo to AMD devices [1], and have trained GPT2 from scratch with 10B tokens of fineweb-edu on a 4x 7900XTX machine in just a few hours (about $2 worth of electricity) [2]. I've also trained the larger GP
10.
▲
by
anthonix1
2y ago
Hi, why do you believe that bfloat16 is not supported? Can you please provide some references (specifically the part about the hardware "doesn't do it")? For the hardware you are focussing on (gfx11), the reference manual [2]
11.
▲
by
anthonix1
2y ago
I just tried it with llm.c ... seems to be missing quite a few key components such as cublaslt, bfloat16 support, nvtx3, compiler flags such as -t And its linked against an old release of ROCm. So unclear to me how it is supposed to be an i
12.
▲
by
anthonix1
2y ago
I have not been impressed by the perf. Slower than PyTorch for LLMs, and PyTorch is actually stable on AMD (I've trained 7B/13B models).. so the stability issues seem to be more of a tinygrad problem and less of an AMD problem, de
13.
▲
by
anthonix1
2y ago
Final loss from that fineweb-10B run (since then I'm up to ~100k toks/sec/GPU): step 18865/18865 | train loss 3.280550 | norm 0.4362 | lr 0.00e+00 | 1669.06 ms | 55.4% A100 fp16 MFU | 314058 tok/s Writing state to l
14.
▲
by
anthonix1
2y ago
55.4% in the last run, at running temperature
15.
▲
by
anthonix1
2y ago
Yeah, I just reproduced the GPT2 from scratch results in 8.75 hours on 4x 7900 XTX. The fork is here: https://github.com/anthonix/llm.c
16.
▲
by
anthonix1
2y ago
Ran tinygrad again about a week ago, no change. And still no comment on the issue, will re-run if there is any comment.
17.
▲
by
anthonix1
2y ago
I think the matmul issue is symptomatic of a much deeper issue. It would be nice to see less whining and blaming AMD (PyTorch and llm.c actually work on 7900 XTX, and blow tiny grad out of the water in terms of perf!), and more just getting
18.
▲
by
anthonix1
2y ago
Maybe get a 7900 XTX. 122 TFLOPS of BF16/FP16 for less than $1k and I'm getting 55.4% MFU
19.
▲
by
anthonix1
2y ago
Nah, I reproduced on 4x 7900 XTX machine in 8.75 hours, so a single 7900 XTX (costs less than $1k) could do it in under 24 hours. Was hitting 55.4% MFU.
20.
▲
by
anthonix1
2y ago
So... successfully reproduced in ~8.75 hours, taking about 18 kWh / $2.70 The first run actually failed at step 3000 or so, and I realized I had a bug in my attention / matmul kernels, but after fixing that and restarting it worke
21.
▲
by
anthonix1
2y ago
Even without hipBLASlt, PyTorch is still ~4x faster than tinygrad on a 7900 XTX for GPT2, and works fine. Any idea why?
22.
▲
by
anthonix1
2y ago
Seems to be an issue on their side. E.g., for a step of GPT2 training on a 7900 XTX [1]: tinygrad is ~440ms, PyTorch 2.4.0.dev20240513 is ~97ms, Karpathy's llm.c with ROCm is ~79ms, and llm.c with custom kernels is ~58ms [1] https:&#x
23.
▲
by
anthonix1
2y ago
It converges similarly on smaller datasets. About to kick off a training from scratch run on the same fineweb-10B, which at 324k toks/sec should take about 8.6 hours. And with my kWh cost, that is about $2.50 cost to train. Will report
24.
▲
by
anthonix1
2y ago
FWIW, I'm seeing ~318,000 toks/sec throughput on a 4x AMD 7900 XTX machine (less than $4k worth of GPU), using the same settings as in the post (0.5M batch size etc).