3 ms·
why not just compare against a Pytorch reference instead of comparing the kernels to a separate scalar Rust implementation? (is it for perf reasons because bot
by mfog 17d ago
why not just compare against a Pytorch reference instead of comparing the kernels to a separate scalar Rust implementation?
(is it for perf reasons because bottleneck becomes verification? Though in that case perhaps it's better to just run a Pytorch reference on the GPUs, which should be good enough. Though, perhaps if you could get this current design to the point where `overlap(CPU, GPU cuda) < GPU cuda + GPU reference` that would be pretty awesome.