5 ms·
The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'l
by bdlowery 19d ago
The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark.
Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change.
EDIT: I was using gemini cli... it's not a harness issue lol
- thereitgoes456 19d agoWhy so brazenly confident? Isn’t it possible that the benchmark is correct, and your experience is correct too, but you haven’t tried all the thousand different modalities of work that programming encompasses and so maybe you don’t actually have standing to judge?
- 0x457 19d agovery outdated experience from me: when I first tried gemini something, in an existing rust codebase, it looked around for files that would indicate if its go, javascript, java or c++ project, then declared I must have asked it build a new app in javascript and proceeded to circle around to figure out how it can install node and npm on my machine. So I totally believe that Gemini is just bad. Which is surprising because Gemma is very good for some tasks, but I never ever had any success with Gemini, be it in cli or chat thing or anything else that has gemini branding.
- deleted 19d ago[deleted]
- siddbudd 19d agohavent tried that model, but it sounds like a potential harness issue. Have you tried it in different harnesses?
- tucnak 19d agoHard disagree. I use 3.8 flash in Antigravity a lot, and thoroughly prefer it to most Pro-class models. It's really fast, and I've had it make crazy progress on compiler-like problems that previous models including Opus simply failed at. On ultra plan you can have it going for hours, and make incremental progress with good prompting for review interrupts. It solved a problem I couldn't solve for weeks in under 6 hours. 10k LOC total. The harness and test suite is key.
- bdlowery 19d agoThis is my exact experience with the model - https://x.com/ThePrimeagen/status/2095565354726502683 https://x.com/ThePrimeagen/status/2095565354726502683 And it just BURNS tokens like crazy.
- TomGarden 19d agoYou sure it was 3.8 Flash? It hasn't been called Gemini cli in a WHILE...
- starchild3001 19d agoThere's no such things as gemini cli these days. It's called "agy" (short for antigravity). And if you don't know what that is, you're probably 3-6 months behind already. PS: Just Googled it to confirm: Gemini CLI was deprecated on May 19th, 2026. The correct harness is called agy or antigravity for Gemini 3.8 Flash. https://developers.googleblog.com/an-important-update-transitioning-gemini-cli-to-antigravity-cli/ https://developers.googleblog.com/an-important-update-transi...
- deleted 19d ago[deleted]
- deleted 19d ago[deleted]
- astrostl 19d agoThe benchmark page itself asserts that it used Gemini CLI as a harness. I came to the comments just because I noticed the error. For my part — using agy — I found Gemini 3.8 Flash mid.
- janaksunil 19d agoi appreciate the feedback, the benchmark is primarily long horizon real world engineering tasks on big private codebases.
- evalmaster123 16d agoso 10 minutes is long horizon now?