5 ms·
You can also see this difference in open router. But why is there only thinking flash now?
by k8sToGo 1y ago
You can also see this difference in open router.
But why is there only thinking flash now?
- hnuser123456 1y agoApparently, you can make a request to 2.5 flash to not use thinking, but it will still sometimes do it anyways, this has been an issue for months, and hasn't been fixed by model updates: https://github.com/google-gemini/cookbook/issues/722 https://github.com/google-gemini/cookbook/issues/722
- Tiberium 1y agoIt might be a bit confusing, but there's no "only thinking flash" - it's a single model, and you can turn off thinking if you set thinking budget to 0 in the API request. Previously 2.5 Flash Preview was much cheaper with the thinking budget set to 0, now the price is the same. Of course, with thinking enabled the model will still use far more output tokens than the non-thinking mode.
- davedx 1y agoInteresting design choice, and makes me think of "Thinking, Fast and Slow" by Kahneman. (I thought of it quickly, not slowly, so the comparison may only be surface deep.)