5 ms·
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets... 3.7 used 64M on high: https://arti
by eis 24d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash https://artificialanalysis.ai/models/gemini-3-7-flash
3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- WASDx 24d ago3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.
- zuzululu 24d agoi find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium