3 ms·
As it happens, I just got my employer's permission to release as open source a Triton back-end for Metal and D3D12 GPUs here: https://github.com/dropbox/neso ht
by jacobgorm 17d ago
As it happens, I just got my employer's permission to release as open source a Triton back-end for Metal and D3D12 GPUs here: https://github.com/dropbox/neso https://github.com/dropbox/neso .
As an example of you how can use it to deploy real models there is this project doing ASR and TTS: https://github.com/dropbox/nspeech https://github.com/dropbox/nspeech .
Finally, I am also going to be switching the inferencing part of Witchcraft from current Candle on MacOS and OpenVINO on Windows to just Candle with Neso; https://github.com/dropbox/witchcraft https://github.com/dropbox/witchcraft
- bbkane 17d agoThat sounds like it'll be easier to maintain. Will it also be faster?
- jacobgorm 17d agoIt is currently faster than the stock Candle / MPS shaders it replaces on MacOS/ARM64, and IIRC a bit slower than OpenVINO/CPU on my old Windows laptop, where I never got OpenVINO/GPU to compute correctly. Candle didn't have support for GPUs on MacOS/Intel, and OpenVINO ceased to be supported there. Compared to OpenVINO (I tried ONNX runtime too, but never got it produce correct outputs with my quantized models) it is very nice to be able to build the exact kernels I need, at the quantization settings and precision that works for the models I have and with the custom operators required (speech models do a lot of non-standard stuff), run from a single set of sources, and not have to ship a hefty third-party DLL, and having to deal with their memory leaks and other stability issues.