3 ms·
Ggml / llama.cpp has a lot of hardware optimizations built in now, CPU, GPU and specific instruction sets like for apple silicon (I'm not familiar with the name
by version_five 3y ago
Ggml / llama.cpp has a lot of hardware optimizations built in now, CPU, GPU and specific instruction sets like for apple silicon (I'm not familiar with the names). I would want to know how many of those are also present in onnx and available to this model.
There are currently also more quantization options available as mentioned. Though those incur a performance loss (they make the model faster but worse) so it depends on what you're optimizing for.
- brucethemoose2 3y agoONNX is a format. There are different runtimes for different devices... But I can't speak for any of them. > specific instruction sets like for apple silicon You are thinking of the Accelerate framework support, which is basically Apple's ARM CPU SIMD library. But Llama.cpp also has a Metal GPU backend, which is the defacto backend for Apple devices now.
- deleted 3y ago[deleted]