4 ms·
Is this is what you are talking about? https://www.reddit.com/r/LocalLLaMA/comments/1j1nen4/llms_like_gpt4o_outputs/ https://www.reddit.com/r/LocalLLaMA/comment
by dimitri-vs 2y ago
Is this is what you are talking about? https://www.reddit.com/r/LocalLLaMA/comments/1j1nen4/llms_like_gpt4o_outputs/ https://www.reddit.com/r/LocalLLaMA/comments/1j1nen4/llms_li...
The prompt response they are judging is "Write one concise paragraph about the company that created you" which is kind of an odd choice.
Sonnet3.7 hates its own meta-analysis, but loves gpt4o. But the reason behind that is because Claude 3.7 Sonnet consistently replies (to the prompt) that it was created by Open AI, but then catches itself (when judging) as being wrong on that.
My takeaway was that gpt4o's gave very safe/luke-warm scores and was strongly correlated with sonnet scores. So if you are judging anything using LLM's then taking the average of Anthropic + Gemini or OpenAI + Gemini might be the best approach.
- moffkalast 2y agoAh yeah that's the one, missed that bit entirely. I guess they should've honestly asked them to write an article about cats or something unrelated that wouldn't hit so close to home.