5 ms·
Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly low
by WinstonSmith84 6mo ago
Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.
- muyuu 6mo agoWhen did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.
- gck1 6mo agoTried Gemini 2 weeks ago to see where it's at, with gemini-cli. Failed to use tools, failed to follow instructions, and then went into deranged loop mode. Essentially, it's where it was 1.5 years ago when I tried it the last time. It's honestly unbelievable how Google managed to fail so miserably at this.
- 4b11b4 6mo agoTheir harness might be behind
- gck1 6mo agoI think failures that I observed with gemini are unrelated to the harness. Because the same failures happened with third party harnesses too.
- unsupp0rted 6mo agoIt’s great on AI Studio. Harness issues, I agree.
- Kailhus 6mo agoI have not experienced any issues with Gemini 3.1 Pro.
- sandos 6mo agoAgree, I never actually had great success with Opus. I think its the failures that are annoying, its probably better than codex when its "good", but it fails in annoying ways that I think codex very seldom does.
- StingyJelly 6mo agoI wouldn't call codex considerably better. It may depend on specific codebase and your expectations, but codex produces more "abstraction for the sake of abstraction" even on simple tasks, while opus in my experience usually chooses right level of abstraction for given task.
- otabdeveloper4 6mo agoA secret art known to the cognoscenti as "benchmark gaming".