3 ms·
From the abstract: > (In the common case that one matrix is known ahead of time,) our method also has the in- teresting property that it requires zero multiply
by mandarax8 4y ago
From the abstract:
> (In the common case that one matrix
is known ahead of time,) our method also has the in-
teresting property that it requires zero multiply-adds.
These results suggest that a mixture of hashing, aver-
aging, and byte shuffling—–the core operations of our
method—–could be a more promising building block
for machine learning than the sparsified, factorized,
and/or scalar quantized matrix products that have re-
cently been the focus of substantial research and hard-
ware investment.`
This is not at all what modern gpus are optimized for.