3 ms·
I have seen that kind of drift in smaller models (e.g. DeepSeek V4 Flash) when setting the thinking too high. So less thinking would lead to better results. But
by vincent_s 2mo ago
I have seen that kind of drift in smaller models (e.g. DeepSeek V4 Flash) when setting the thinking too high. So less thinking would lead to better results. But that's not something I'd expect from a SOTA model. Higher thinking effort should lead to same or better results.