3 ms·
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory
by armcat 1mo ago
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory use. It's more about what happens in relation to GGUF, MLX and non CUDA-native frameworks/formats.
- belzebub 1mo agoBingo.