3 ms·
It’s approximately the same as Qwen3.827b’s propensity to think a lot, right?
by nojs 29d ago
It’s approximately the same as Qwen3.827b’s propensity to think a lot, right?
- digdugdirk 29d agoNot quite. Looped models do the extra "thinking" inside the model's layers. So the token gets twice the number crunching performed on it before it gets spit out. I think of it as the first loop "kickstarts" the process, and the second loop refines it.