4 ms·
https://github.com/t0rsion/leone https://github.com/t0rsion/leone Working on my local LLM runtime Leone. It has a CPU path that verifies the CUDA kernels.
by torsi0n 17d ago
https://github.com/t0rsion/leone https://github.com/t0rsion/leone
Working on my local LLM runtime Leone. It has a CPU path that verifies the CUDA kernels.
- mfog 17d agowhy not just compare against a Pytorch reference instead of comparing the kernels to a separate scalar Rust implementation? (is it for perf reasons because bottleneck becomes verification? Though in that case perhaps it's better to just run a Pytorch reference on the GPUs, which should be good enough. Though, perhaps if you could get this current design to the point where `overlap(CPU, GPU cuda) < GPU cuda + GPU reference` that would be pretty awesome.