4 ms·
Just ran it on one of our internal PDF (3 pages, medium difficulty) to json benchmarks: gemini-flash-2.0: 60 ish% accuracy 6,250 pages per dollar gemini-2.5-
by serjester 1y ago
Just ran it on one of our internal PDF (3 pages, medium difficulty) to json benchmarks:
gemini-flash-2.0:
60 ish% accuracy
6,250 pages per dollar
gemini-2.5-flash-preview (no thinking):
80 ish% accuracy
1,700 pages per dollar
gemini-2.5-flash-preview (with thinking):
80 ish% accuracy (not sure what's going on here)
350 pages per dollar
gemini-flash-2.5:
90 ish% accuracy
150 pages per dollar
I do wish they separated the thinking variant from the regular one - it's incredibly confusing when a model parameter dramatically impacts pricing.
- ValveFan6969 1y agoI have been having similar performance issues, I believe they intentionally made a worse model (Gemini 2.5) to get more money out of you. However, there is a way where you can make money off of Gemini 2.5. If you set the thinking parameter lower and lower, you can make the model spew absolute nonsense for the first response. It costs 10 cents per input / output, and sometimes you get a response that was just so bad your clients will ask for more and more corrections.
- mpalmer 1y agoWow, what apps have you made so I know never to use them?