5 ms·
The motivation for AutoThink came from watching how current reasoning models waste computation - they spend the same amount of "thinking time" on "what's 2+2?"
by codelion 1y ago
The motivation for AutoThink came from watching how current reasoning models waste computation - they spend the same amount of "thinking time" on "what's 2+2?" as they do on complex mathematical proofs. This seemed obviously inefficient.
The breakthrough was combining two techniques I'd been working on separately: adaptive classification (which can learn new categories without retraining) and an open source implementation of Pivotal Token Search from Microsoft's Phi-4 paper. When I put them together with dynamic token budgeting, the performance gains were much better than expected.
What surprised me most was that the technique actually uses fewer tokens on average while improving performance. The adaptive allocation means simple queries finish faster, offsetting the extra computation on complex ones.
A few technical notes:
- The steering vectors are small (typically <1MB per pattern) and add minimal memory overhead
- Classification adds about 10ms latency, which is negligible
- Target layer selection matters - I found middle layers (15-20) work best for most models
I'd love feedback on:
- Have you tried similar adaptive approaches with your models?
- What other reasoning patterns would be useful to steer toward?
- Ideas for automatically detecting the optimal target layer?
Thanks for checking it out! Happy to answer any questions about the implementation or results.
- behnamoh 1y ago> they spend the same amount of "thinking time" on "what's 2+2?" as they do on complex mathematical proofs. Not anymore. Have you seen Gemini 2.5 Pro? Ask it simple questions and it almost doesn't "think". Ask it a coding question and it'll write a long reasoning article. I think the same goes for o3.
- codelion 1y agoYes, we started with the idea of trying to replicate similar control on thinking processes for open reasoning models. They also announced the Deep Think approach at IO which goes even further and combines parallel CoTs at inference.
- sigmoid10 1y agoThe original o1 also didn't do this. Neither did the actual DeepSeek R1. You could even get it to answer immediately without any reasoning tokens. These highly distilled versions just lost most of their common sense for this.
- shing3232 1y agoWell, it does overthink quite a bit. if It can reduce overthink,it s gonna be useful
- victorbjorklund 1y agoOverthink is subjectibe. It really depends on how much you value the answer. "how long break distance does a train need if going in 100 km/hour?" Just need a quick reply and you dont care so much (maybe showerthought)? Or is life and death depending on the answer? The same question can need different amount of thinking.
- normie3000 1y ago> is life and death depending on the answer? In this situation I suspect you'd still want the answer quickly.
- diggan 1y agoHuge assumption, there is a wide range of various parameters that goes into how accurate you need an response to be, depending on context. As sure as there exists questions that you need 100% accurate response regardless of response times, I'm sure there exists questions on the other extreme.
- GTP 1y agoIn this situation you would have someone with actual knowledge of the mechanics involved do the computation using the actual data (e.g., what's the mass of the train? Which kind of breaks does it have?) instead of asking an LLM and trusting it to give the correct answer without checking.
- CjHuber 1y agoWhat I really don't like is that I can't manually decide how much thinking it Gemini should allocate to a prompt. You're right sometimes it doesn't think but for me this also happens on complex query where I WOULD want it to think. Even things like "super think about this" etc don't help, it just refuses to
- thegeomaster 1y agoGemini 2.5 Pro is getting thinking budgets when it GAs in June (at least that's the promise).
- vladf 1y agoThis is available for Flash
- CharlesW 1y ago> I think the same goes for o3. Definitely, in my experience. Elsewhere in the thread, OP says that open models/systems don't do this, in which case this seems like important work toward making open alternatives competitive.
- olddustytrail 1y agoIs that not just caching? If you have the same query just return the same response. You could even put a simpler AI in front to decide if it was effectively the same query.
- mclau157 1y agoHas Gemini or OpenAI put out any articles on this or is this just something you noticed?
- Abishek_Muthian 1y agoCongratulations! Any work to optimise efficiency w.r.t LLMs is much appreciated. So far I’ve taken only lazy approach to optimising local LLMs by sending small queries to my M4 Mac Mini running MLX models and sending larger queries to my Nvidia 4090; it’s remarkable how efficient M4 is compared to Nvidia and I think Apple is in the right direction with MLX. I would read about AutoThink and try to integrate it with my workflow.
- Lerc 1y agoI have thought it might be worth seeding responses with the output of non-reasoning models, so after the user prompt, inject a block of "a non-reasoning model thought this:... stuff ....Was that what the user wanted?" For the instances where the non reasoning version was sufficient it might help the reasoning model get to the point earlier.
- codelion 1y agoThis is an interesting idea, I hadn't thought of it. It is worth experimenting I am not aware of anyone else trying it yet.
- waffletower 1y agoClaude Sonnet 3.5 (not even the latest iterations: 3.7 or 4) clearly adapts processing time to query complexity -- processing time is dynamic.