3 ms·
onnxruntime is typically for GPU acceleration while being nearly unusable on CPUs. It’s also better supported (Microsoft) and supports LOTS of APIs. Llama.cpp
by Jaxkr 3y ago
onnxruntime is typically for GPU acceleration while being nearly unusable on CPUs. It’s also better supported (Microsoft) and supports LOTS of APIs.
Llama.cpp / ggml (while they support some hardware acceleration) is more focused on commodity hardware like x86 CPUs and Apple M-series silicon.
- sroussey 3y agoHardware acceleration includes: OpenBLAS/Apple BLAS/ARM Performance Lib/ATLAS/BLIS/Intel MKL/NVHPC/ACML/SCSL/SGIMATH and more in BLAS. Also new apple metal implementation in progress in addition to apple accelerate, and is in baseline if you enable it. Also a CUDA implementation.