9 ms·
Beats Gemini 1.5 Pro at all but two of the listed benchmarks. Google DeepMind is starting to get their bearings in the LLM era. These are the minds behind Alpha
by bradhilton 2y ago
Beats Gemini 1.5 Pro at all but two of the listed benchmarks. Google DeepMind is starting to get their bearings in the LLM era. These are the minds behind AlphaGo/Zero/Fold. They control their own hardware destiny with TPUs. Bullish.
- p1esk 2y agoAre these benchmarks still meaningful?
- maeil 2y agoNo, and they haven't been for at least half a year. Utterly optimized for by the providers. Nowadays if a model would be SotA for general use but not #1 on any of these benchmarks, I doubt they'd even release it.
- CamperBob2 2y agoI've started keeping an eye out for original brainteasers, just for that reason. GCHQ's Christmas puzzle just came out [1], and o1-pro got 6 out of 7 of them right. It took about 20 minutes in total. I wasn't going to bother trying those because I was pretty sure it wouldn't get any of them, but decided to give it an easy one (#4) and was impressed at the CoT. Meanwhile, Google's newest 2.0 Flash model went 0 for 7. 1: https://metro.co.uk/2024/12/11/gchq-christmas-puzzle-2024-released-can-solve-22161098 https://metro.co.uk/2024/12/11/gchq-christmas-puzzle-2024-re...
- nrvn 2y agoDid it get the 8 right? The linked article provides the wrong answer btw.
- CamperBob2 2y agoI didn't see a straightforward way to submit the final problem, because I used different contexts for each of the 7 subproblems. Given the right prompt, though, I'm sure it could handle the 'find the corresponding letter from the landmarks to form an anagram' part. That's easier than most of the other problems. You're saying the ultimate answer isn't 'PROTECTING THE UNITED KINGDOM'?
- nrvn 2y agoif you follow the sleigh morse path starting from the robin it will be 'united in protecting the kingdom'.
- p1esk 2y agoWow! That’s all I need to know about Google’s model.
- danpalmer 2y agoThat's a comparison of multiple GPT-4 models working together... against a single GPT-4 mini style model.
- p1esk 2y agomultiple GPT-4 models working together What do you mean? Is o1 not a single model?
- Workaccount2 2y agoWhat is impressive about this new model is that it is the lightweight version (flash). There will probably be a 2.0 pro (which will be 4o/sonnet class) and maybe an ultra (o1(?)/Opus).
- iamdelirium 2y agoWhy are you comparing flash vs o1-pro, wouldn't a more fair comparison be flash vs mini?
- iamdelirium 2y agoI just ask o1-mini the first two questions and it got it wrong.
- deleted 2y ago[deleted]
- CamperBob2 2y agoIt's the only Google model that my account has access to that accepts .PNG files. I assumed it was the latest/greatest experimental 2.0 release. If they want a rematch, they'll need to bring their 'A' game next time, because o1-pro is crazy good.
- dagmx 2y agoRegarding TPU’s, sure for the stuff that’s running on the cloud. However their on device TPUs lag behind the competition and Google still seem to struggle to move significant parts of Gemini to run on device as a result. Of course, Gemini is provided as a subscription service as well so perhaps they’re not incentivized to move things locally. I am curious if they’ll introduce something like Apple’s private cloud compute.
- whimsicalism 2y agoi don’t think they need to win the on device market. we need to separate inference and training - the real winners are those who have the training compute. you can always have other companies help with inference
- dagmx 2y agoAt what point does the on device stuff eat into their market share though? As on device gets better, who will pay for cloud compute? Other than enterprise use. I’m not saying on device will ever truly compete at quality, but I believe it’ll be good enough that most people don’t care to pay for cloud services.
- whimsicalism 2y agoYou're still focused about inference :) inference basically does not matter, it is a commodity
- dagmx 2y agoYou’re still focused about training :) training doesn’t matter if inference costs are high and people don’t pay for them
- whimsicalism 2y agobut inference costs arent high already and there are tons of hardware companies that can do relatively cheap LLM inference
- JeremyNT 2y agoYeah they've been slow to release end-user facing stuff but it's obvious that they're just grinding away internally. They've ceded the fast mover advantage, but with a massive installed base of Android devices, a team of experts who basically created the entire field, a huge hardware presence (that THEY own), massive legal expertise, existing content deals, and a suite of vertically integrated services, I feel like the game is theirs to lose at this point. The only caution is regulation / anti-trust action, but with a Trump administration that seems far less likely.
- VirusNewbie 2y agoIf you look at where talent is going, it's Anthropic that is the real competitor to Google, not OpenAI.