2 ms·
This is all quantifiable. I regularly have my model run benchmarks against all the config permutations and then choose the best based on my criteria, which typi
by Zetaphor 2mo ago
This is all quantifiable. I regularly have my model run benchmarks against all the config permutations and then choose the best based on my criteria, which typically boil down to trading prefill and decode times
- LoganDark 2mo agoWhat model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.
- Zetaphor 2mo agoI'm running on Strix Halo so memory bandwidth is my constraint. In that example I'm describing the choice between using ROCm or Vulkan. I have a llama-swap config that can call different instances of llama-server running a toolbox with either runtime.
- LoganDark 2mo agoMemory bandwidth is my constraint too (M4 Max) but prefill and single-token decode don't run at the same time. It's best to use batched prefill so you can benefit from processing multiple tokens with a single pass through the model weights.