4 ms·
Are you insinuating Gemini is similar in performance to o3-mini?
by BinRoo 2y ago
Are you insinuating Gemini is similar in performance to o3-mini?
- gerdesj 2y agoAre you implying it isn't? (evidence please, everyone)
- BinRoo 2y agoSimple example: o3-mini-high gets this [1] right, whereas Gemini 2.0 Flash 01-21 gets it wrong. [1] https://chatgpt.com/share/679d9579-5bb8-8008-ac4a-38cef65b45b5 https://chatgpt.com/share/679d9579-5bb8-8008-ac4a-38cef65b45...
- xnx 2y agoGreat example. Thank you. Can confirm that none of the Gemini models warned about the exception without prompting.
- maeil 2y agoThis agrees with my limited testing so far, but in a different way: o3 being better at coding and objective tasks, with the most recent Flash 2.0-thinking stronger at subjective tasks. Similarly, o3 seems better at shorter output sizes, but drops off, tending to be lazy.
- xnx 2y agoDefinitely varies by application, but the blind "taste test" vibes are very good for Gemini: https://lmarena.ai/?leaderboard https://lmarena.ai/?leaderboard
- anabab 2y agothat reminds me that a week ago there was a (now deleted but has a copy of the content available in the comments) post on Reddit where the author claimed they have attempted manipulating/manipulated voting on lmarena in favor of Gemini to tip the scale on Polymarket where on a question like "which AI model will be the best one by $date" (with the outcome decided based on the scoring on lmarena) they have supposedly made O(USD10k). Original deleted post: https://old.reddit.com/r/MachineLearning/comments/1i83mhj/lm_arena_public_voting_is_not_objective_for_llm/ https://old.reddit.com/r/MachineLearning/comments/1i83mhj/lm... A copy of the content: https://old.reddit.com/r/MachineLearning/comments/1i83mhj/lm_arena_public_voting_is_not_objective_for_llm/m9029so/ https://old.reddit.com/r/MachineLearning/comments/1i83mhj/lm...
- deleted 2y ago[deleted]
- panarky 2y agoI've only had o3-mini for a day, but Gemini 2.0 Flash Thinking is still clearly better for my use cases. And it's currently free in aistudio.google.com and in the API. And it handles a million tokens.