5 ms·
Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 1
by Buttons840 1y ago
Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today.
(I'm mostly making this comment to document what happened for the history books.)
https://polymarket.com/event/which-company-has-best-ai-model-end-of-2025 https://polymarket.com/event/which-company-has-best-ai-model...
- roflyear 1y agoThe Musk effect is pretty crazy. Or is there another explanation for why x can compete with Google?
- Davidzheng 1y agoThey have a lot of compute already and Grok 4 was pretty strong?
- whimsicalism 1y agothey’ve managed to acquire compute remarkably quickly and i’m no Musk lover
- deltaburnt 1y agoThinking more cynically: political corruption and connections I'm guessing? Just a couple months ago Musk was treating the US government like his personal playground.
- ENGNR 1y agoElon's Y Combinator interview was pretty good. He seemed more in his element back amongst the hacker crowd (rather than dirty politics), and seemed to be doing hackery things at X, like renting generators and mobile cooling vans and just putting them the car park outside a warehouse to train Grok, since there were no data centres available and he was told it would take 2 years to set it all up properly. I think he's just good at attracting good talent, and letting them focus on the right things to move fast initially, while cutting the supporting infra down to zero until it's needed.
- criddell 1y agoAre you talking about this: https://futurism.com/elon-musk-memphis-illegal-generators https://futurism.com/elon-musk-memphis-illegal-generators It's hackery but also kind of sociopathic to dump a bunch of loud, dirty generators in the middle of a low-income community. Go set your data center up on Martha's Vineyard and see how long the residents put up with it.
- raincole 1y agoBecause they started so late but somehow managed to make something close to SOTA? Either way or people think Trump will just give Elon a 500B government contract...
- apetresc 1y agoHow on Earth does that market have Anthropic at 2%, in a dead heat with the likes of Meta? If the market was about yesterday rather than 5 months from now I think Claude would be pretty clearly the front runner. Why does the market so confidently think they’ll drop to dead last in the next little while?
- epiccoleman 1y agoI find this confusing too. I dropped my OpenAI subs for Claude a while back and I don't feel like I'm missing much. I need to spend some more time with Gemini too though. I was using that as a backend for Cursor for a while and had some good results there too.
- manmal 1y agoClaude is a useful tool, IMO the most useful one even, but not a road to AGI.
- Buttons840 1y agoHow is Claude doing on the benchmark that market is based on? Maybe not so good? Idk. Just because Claude is good for real world use doesn't mean it's winning the benchmark, but the benchmark is all that matters for the Polymarket.
- vasco 1y agoIf you think it's wrong, participate. That's the only way prediction markets end up predicting anything.
- Tadpole9181 1y agoAh, yes, if you disagree you must participate in real money gambling based on the outcome of a single user-based, single-prompt leaderboard.
- vasco 1y agoWell I for example don't give a shit what prediction markets do and never participated, but if someone thinks they're wrong, they should just participate and get free money. Otherwise why complain.
- m3kw9 1y agoIs not that they are not impressed, is just google came out with steerable video gen
- Buttons840 1y agoThat was a few days ago. The big drop in that Polymarket I mentioned all happened today. It was reaction to GTP5 specifically.
- jstummbillig 1y agoThat bet does not seem to be very illuminating. Winner is likely who happens to release closest to end of year, no?
- boringg 1y agoYou don't actually hold polymarket odds with any significant weighting on actual outcomes do you?
- vessenes 1y agoAfter a few hours with gpt-5, I'd trade that spread. Not that I think oAI will win end of year. But I think gpt5 is better than it looks on the benchmark side. It is very very good at something we don't have a lot of benchmarks for -- keeping track of where it's at. codex is vassstly better in practice than claude code or gemini cli right now. On the chat side, it's also quite different, and I wouldn't be surprised if people need some time to get a taste and a preference for it. I ask most models to help me build a macbook pro charger in 15th century florence with the instructions that I start with only my laptop and I can only talk for four hours of chat before the battery dies -- 5 was notable in that it thought through a bunch of second order implications of plans and offered some unusual things, including a list of instructions for a foot-treadle-based split ring commutator + generator in 15th century florentine italian(!). I have no way of verifying if the italian was correct. Upshot - I think they did something very special with long context and iterative task management, and I would be surprised if they don't keep improving 5, based on their new branding and marketing plan. That said, to me this is one of the first 'product release' moments in the frontier model space. 5 is not so much a model release as a polished-up, holes-fixed, annoyances-reduced/removed, 10x faster type of product launch. Google (current polymarket favorite) is remarkably bad at those product releases. Back to betting - I bet there's a moment this year where those numbers change 10% in oAIs favor.
- ttroyr 1y agoI would agree. I am a big fan of Claude and I've Claude code a bunch although after testing Codex & GPT-5 extensively, it just gets stuck in a rut way less often and much more often is able to pinpoint issues & fixes in the codebase.
- croemer 1y agoLooking at LMarena which polymarket uses, I'm not surprised. Based on the little data there is (3k duels, it's possibly worse than Gemini, it lost more to Gemini 2.5 Pro than it won in direct duels). Not sure why the ELO is still higher, possibly GPT5 did more clearly better against bad models, which I don't care about.
- riku_iki 1y ago> Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end) who will decide the winner to resolve bets?