3 ms·
74.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Clau
by oof-baroomf 1y ago
74.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Claude 3.5 Sonnet
- Workaccount2 1y agoI'm pretty sure 3.5 sonnet always benchmarked poorly, despite it being the clear programming winner of it's time.
- iLoveOncall 1y agoThat would assume there is a giant breakthrough to be found.