4 ms·
gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contex
by codingwagie 1y ago
gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contextual prompt ( > 50k tokens) with a ton of code, and see how bad it is. I cancelled my gemini subscription
- lerchmo 1y agohttps://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/ your experience doesn't align with my experience or this benchmark. o1 pro is good but I would rather do 20 cycles on gemini 2.5 rather than wait for Pro to return.
- jjani 1y agoI have, dozens of times, and it's generally better than 3.7. Especially with more context it's less forgetful. o1-pro is absurdly expensive and slow, good luck using that with tools. Virtually all benchmarks, including less gamed ones such as Aider's, show the same. WebLM still has 3.7 ahead, with Sonnet always having been particularly strong at web development, but even on there 2.5 Pro is miles in front of any OpenAI model. Gemini subscription? Surely if you're "using LLMs correctly" you'd have been using the APIs for everything anyway. Subscriptions are generally for non-techy consumers. In any case, just straight up saying "it isn't good" is absurd, even if you personally prefer others.
- int_19h 1y agoI have just watched Sonnet 3.7 vs Gemini 2.5 solving the same task (fix a bug end-to-end) side by side, and Sonnet hallucinated far worse and repeatedly got stuck in dead-ends requiring manual rescue. OTOH Gemini understood the problem based on bug description and code from the get go, and required minimal guidance to come up with a decent solution and implement it.