3 ms·
slow as in the output is slow, or slow as in slow token rates? qwen3.8 has been fantastic as far as token rates are concerned for me, but the overthinking thin
by serf 15d ago
slow as in the output is slow, or slow as in slow token rates?
qwen3.8 has been fantastic as far as token rates are concerned for me, but the overthinking thing with higher reasoning levels takes some coercion to get right.
fwiw pi and hermes both handle that model fairly well. omp required a lot of tuning. I didn't bother figuring out why, I presume it's because qwen3.8 expects reasoning declarations a bit differently. nothing a proxy can't fix.