3 ms·
that’s been wrong for a while, but affirmed with gpt-4o-mini beating out Sonnet 3.5. OpenAI fine tuned 4o and 4o-mini to provide answers that meaningfully impro
by laborcontract 2y ago
that’s been wrong for a while, but affirmed with gpt-4o-mini beating out Sonnet 3.5. OpenAI fine tuned 4o and 4o-mini to provide answers that meaningfully improve model congeniality but trivially improve model intelligence.
Chatbot Arena ELO is a dead metric.
- thomasahle 2y agoWow, I overlooked GPT-4o-mini that far up. But if you change the category to Math (or something else hard), mini drops way down and Claude 3.5 Sonnet goes to the top.