5 ms·
What are the implications of onnxruntime vs something like llama.cpp or those nn-512 models that seem closer to just C code for inference? What's different?
by version_five 3y ago
What are the implications of onnxruntime vs something like llama.cpp or those nn-512 models that seem closer to just C code for inference? What's different?
- deleted 3y ago[deleted]
- Jaxkr 3y agoonnxruntime is typically for GPU acceleration while being nearly unusable on CPUs. It’s also better supported (Microsoft) and supports LOTS of APIs. Llama.cpp / ggml (while they support some hardware acceleration) is more focused on commodity hardware like x86 CPUs and Apple M-series silicon.
- sroussey 3y agoHardware acceleration includes: OpenBLAS/Apple BLAS/ARM Performance Lib/ATLAS/BLIS/Intel MKL/NVHPC/ACML/SCSL/SGIMATH and more in BLAS. Also new apple metal implementation in progress in addition to apple accelerate, and is in baseline if you enable it. Also a CUDA implementation.
- jxy 3y agoAny POSIX system that has a C++11 compiler could compile llama.cpp and run it, including all the BSDs. The OS supported by onnxruntime is very limited.
- IshKebab 3y agoIt looks like onnxruntime supports all major OSes and the web.