4 ms·
It is very likely that they compared newly supported BF16 instruction set with some other (float? fp16? quantized ints?). If you make simple 1-to-1 transition
by Lockal 3y ago
It is very likely that they compared newly supported BF16 instruction set with some other (float? fp16? quantized ints?).
If you make simple 1-to-1 transition from AVX256 to AVX512, the speedup is usually <2x, unless there is a major bottleneck in instruction fetch and decoding. Also AVX512 in Zen is still double-pumped. Regarding FMA, again, if you compare FMA in AVX256 and AVX512, latencies and throughput are the same[1].
Comparing performance between different datatypes is probably fine, but it should be stated directly. Unless "Owners of CPUs like Zen4 can expect to see 10x faster prompt eval times" means comparison with Skylake.
[1] https://uops.info/table.html?search=vfmadd132ps&cb_lat=on&cb_tp=on&cb_uops=on&cb_ports=on&cb_ZEN4=on&cb_measurements=on&cb_doc=on&cb_avx512=on&cb_fma=on https://uops.info/table.html?search=vfmadd132ps&cb_lat=on&cb...
- janwas 3y agoNice result, congrats Justine! The bf16 dot instruction replaces 6 instructions: https://github.com/google/highway/blob/master/hwy/ops/x86_128-inl.h#L9214 https://github.com/google/highway/blob/master/hwy/ops/x86_12... A 3-4x speedup vs SKX sounds quite plausible :)