3 ms·
What is the purpose of this? Just a hard cutoff below the actual context window? You could set that in your harness anyway.
by hendersoon 2mo ago
What is the purpose of this? Just a hard cutoff below the actual context window? You could set that in your harness anyway.
- surgical_fire 2mo ago> k3 (1M) consumes about twice as much quota as k3-256k Cheaper?
- chrisweekly 2mo agok3-256 consumes half the quota, so... yes? (assuming the user isn't making use of the larger context window)
- dnlzro 2mo agoUses less quota (i.e., cheaper). For people who like to keep their contexts small, this is a no-brainer.
- meatmanek 2mo agoAt least in the self-hosted LLM inference engines, you have to pre-allocate space for the maximum amount of context you want to allow for each parallel session. By using a lower maximum, you don't have to allocate as much VRAM for each session, allowing more usage for the same amount of hardware. Thus, cheaper.