3 ms·
> (Optimized by you through testing. Not that AI) Why not optimized by AI through testing ? Give it a test set to work on and let it loose.
by mrighele 2mo ago
> (Optimized by you through testing. Not that AI)
Why not optimized by AI through testing ? Give it a test set to work on and let it loose.
- LoganDark 2mo agoAI doesn't necessarily know what feels like a good tradeoff to you. I'm sure it could help guide you though.
- hypfer 2mo ago^ This. Intent is the answer and AI has none.
- Zetaphor 2mo agoThis is all quantifiable. I regularly have my model run benchmarks against all the config permutations and then choose the best based on my criteria, which typically boil down to trading prefill and decode times
- LoganDark 2mo agoWhat model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.
- Zetaphor 2mo agoI'm running on Strix Halo so memory bandwidth is my constraint. In that example I'm describing the choice between using ROCm or Vulkan. I have a llama-swap config that can call different instances of llama-server running a toolbox with either runtime.
- LoganDark 2mo agoMemory bandwidth is my constraint too (M4 Max) but prefill and single-token decode don't run at the same time. It's best to use batched prefill so you can benefit from processing multiple tokens with a single pass through the model weights.