5 ms·
Ran some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all.
by _2d30 6mo ago
Ran some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all.
Major analytical errors in their response to multiple of my technical questions.
- _2d30 6mo agoPlaying with this some more and it's actively not good. Just basic mathematical errors riddling responses. Did some basic adversarial testing where its responses are analyzed by Gemini and Gemini is finding basic math errors across every relatively (relative to Opus, Gemini or GPT can handle) simple ask I make. Yikes.
- smlacy 6mo agoPost actual results, make a blog post. Don't just say "this sucks" without tangible evidence. Otherwise you're doomed to "sample size of one" level of relevance.
- thorum 6mo agoI have the opposite experience: random HN/Reddit comments saying “this sucks” or “whoa this is a huge improvement” are the only benchmark that means anything. Standard benchmarks are all gamed and don’t capture the complexity of the real world.
- titanomachy 6mo agoThen your internal benchmarks will be in the post-training set and you’ll have to make new ones.
- _2d30 6mo agoI may already have but I'm pseudonymous on this website.
- smlacy 6mo ago[flagged]
- mliker 6mo agoIt’s quite good for multimodal cases that 3 billion people would use it for though it lags in scientific areas
- _2d30 6mo agoYes, this would make sense for what Meta might focus on.
- jatora 6mo agoeven gemini is not in that conversation