3 ms·
What harness would you recommend? I’ve tried Pi but the model struggled to stay on track after the compaction. I have only 48gb of ram, so can fit only 80k con
by mrtsepelev 1mo ago
What harness would you recommend? I’ve tried Pi but the model struggled to stay on track after the compaction.
I have only 48gb of ram, so can fit only 80k context max, so good compaction is must.
- Scaled 1mo agoNot op, but check out open code; you can turn on K/V quantization to help with increasing context if you have not already. I think K needs to stay at least 8 but I hear V can go down to 4?
- npodbielski 1mo agoI am running it on 32GB and I did not saw model loosing it context even after 4-5 compactions in pi. I am running sessions for few days sometimes. I think it looped once, but loop police extension stopped it. The only problem I have know is how pi compaction works, which is forcing full prefill which takes time and it is erroring a lot. I wrote my own compaction that should remove full prefil but it does not work. But this is the only problem with this setup and it is more problem with pi then the model. I much more prefer it to use Qwen then paid models: Claude forces me to do reauth every other day and codex models either are too costly or not capable enough.