3 ms·
How does this compare to using https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp with https://huggingface.co/models?search=thebloke/
by rrherr 3y ago
How does this compare to using https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp with https://huggingface.co/models?search=thebloke/llama-2-ggml https://huggingface.co/models?search=thebloke/llama-2-ggml ?
- moffkalast 3y agoThese are still FP16/32 models, almost certainly a few times slower and larger than the latest N bit quantized GGMLs.
- version_five 3y agoGgml / llama.cpp has a lot of hardware optimizations built in now, CPU, GPU and specific instruction sets like for apple silicon (I'm not familiar with the names). I would want to know how many of those are also present in onnx and available to this model. There are currently also more quantization options available as mentioned. Though those incur a performance loss (they make the model faster but worse) so it depends on what you're optimizing for.
- brucethemoose2 3y agoONNX is a format. There are different runtimes for different devices... But I can't speak for any of them. > specific instruction sets like for apple silicon You are thinking of the Accelerate framework support, which is basically Apple's ARM CPU SIMD library. But Llama.cpp also has a Metal GPU backend, which is the defacto backend for Apple devices now.
- deleted 3y ago[deleted]
- brucethemoose2 3y agoVery unfavorably. Mostly because the ONNX models are FP32/FP16 (so ~3-4x the RAM use), but also because llama.cpp is well optimized with many features (like prompt caching, grammar, device splitting, context extending, cfg...) MLC's Apache TVM implementation is also excellent. The autotuning in particular is like black magic.
- skeletoncrew 3y agoI tried quite a few of these and the ONNX one seems the most elegantly put together of all. I’m impressed. Speed can be improved. Quick and dirty/hype solutions, not sure. I really hope ONNX gets traction it deserves.
- brucethemoose2 3y ago> ONNX one seems the most elegantly put together of all. What do you mean by this? The demo UI? Code quality?
- version_five 3y ago> Quick and dirty/hype solutions, not sure. Curious what you mean by this
- refulgentis 3y agoIt's tough to hear and communicate, but TL;DR: it's very good for HN headlines to make a $X.cpp but it's not the right tool for products. There's only going to be more of these models and ONNX starts from the right place, cross-platform and from base principles rather than coupling tightly to one model structure. Most importantly, it is freakin' awesome, the comments thus far, 30 in, don't reflect what its like to use or it's technical realities.* * the main threads of discussion are "not even wrong": float16 is big compared to float4 (its trivial to quantize to your liking) and looking for an alternative that supports CoreML (ONNX is the magic that makes your model take advantage of CoreML / WebGPU / WebGL / whatever Android's marketing name for its API is etc. etc. etc.)
- kiratp 3y agoWhen hardware is so expensive and so difficult to obtain, performance starts to trump the cost of “single model implementation” pretty quickly. It can be cheaper to deploy Llama.cpp and then foobar.cpp (6 months from now) than it is to have inference that is 2x slower. Interestingly in the LLM space all these model servers seem to be converging to using the same API as OpenAI, making it easy to swap containers to get a different model+inference server with 0 code change.