2 ms·
Indeed it could be a target, but newer and newer co-processor for pure MatMul or even NPU are created and cache/register are changing fast from one architecture
by machinekob 3y ago
Indeed it could be a target, but newer and newer co-processor for pure MatMul or even NPU are created and cache/register are changing fast from one architecture to another so it still wouldnt be optimal for most users.