3 ms·
Interesting! What's the prompt eval processing speed like compared to llama.cpp and kin?
by castles 3y ago
Interesting! What's the prompt eval processing speed like compared to llama.cpp and kin?
- woadwarrior01 3y agoI haven't run any specific low level benchmarks, lately. But chunked prefilling and tvm auto-tuned Metal kernels from mlc-llm seemed to make a big differenced, the last time I checked. Also, compared to stock mlc-llm, I use a newer version of metal (3.0) and have a few modifications to make models have a slightly smaller memory and disk footprint, also slightly faster execution. Because unlike the mlc-llm folks, I only care about compatibility with Apple platforms. They support so much more than that in their upstream project.
- castles 3y agothanks, I'll give it a crack