3 ms·
I did a similar optimization via https://github.com/viterin/vek https://github.com/viterin/vek as the SIMD version. Some somewhat unscientific calculations show
by huac 3y ago
I did a similar optimization via https://github.com/viterin/vek https://github.com/viterin/vek as the SIMD version. Some somewhat unscientific calculations showed a 10x improvement staying in float32: https://github.com/stillmatic/gollum/blob/07a9aa35d2517af8cfa36d0ee61736010ad91b58/math_test.go#L124 https://github.com/stillmatic/gollum/blob/07a9aa35d2517af8cf... (comparable to the 9x improvement in article using SIMD + int8)
TBH my takeaway was that it was more useful to use smaller vectors as a representation