2 ms·
We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude
by satvikpendem 24d ago
We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.
- onlyrealcuzzo 24d agoAnd the benchmarks agreed with you... until now. So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.
- NitpickLawyer 24d agoIf anything, gemini models are the least benchmaxxed out of any lab, IMO.