4 ms·
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wou
by HAL3000 29d ago
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
- atonse 29d agoyeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall). Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
- elAhmo 29d agoCould you share more about 5x/20x? I missed that
- m101 29d ago20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number. OpenAI is 20x on both limits
- beydogan 29d ago> Weekly seems to be around 10x Actually no. 5x and 20x have same weekly usage across all models. Just ask their chatbot. https://x.com/beydogan_/status/2095293596198957418 https://x.com/beydogan_/status/2095293596198957418
- chid 28d agoit's clearly wrong, think it's realistically closer to 1.7x
- elAhmo 27d agoIs this official or based on people's reports?
- aizk 28d agoI still use Opus 4.8 for a lot of tasks because I can't stand the way it talks.
- CSMastermind 29d agoSol easily outperforms Fable on every task I've tried it on.
- enraged_camel 29d agoI can't speak for others but I have a feeling you're in the very small minority with this take. You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.
- athrowaway3z 28d agoI dont think that feeling is entirely useful. Because Claude doesn't allow third party harnesses on their subscriptions I doubt the majority of signals you're getting are actually that significant on pure model quality. I suspect you're right on Sol not outperforming Fable; but i've not used Fable that much. --- But, fwiw, in my custom harness between Sol & Opus 4.8 - then Sol wins by a ridiculous margin as Opus keeps claiming slightly wrong things with certainty much more.
- andxor 28d agoThis is not saying much. Opus 4.8 is ancient history.
- andxor 29d agoThat's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?
- Koffiepoeder 28d agoFor me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer. Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.
- crossroadsguy 28d ago> Canceling my Anthropic Max sub when this ships. At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.
- iknowstuff 28d agoYeah I have a main $100 sub, a bunch of $20 subs and sometimes another simultaneous $100 when a really major model happens to drop. For the most part it seems better to have multiple subscriptions than a single $200-$300 one to stay more in touch with state of the art and get a feel for what's good at what.
- mlinsey 28d agoEspecially with these big models chewing up limits fast, I feel good about simultaneously having an Anthropic sub, an OpenAI sub, and an Opencode balance. The models also seem to catch things when code reviewing each other that they don't always catch when a new instance of the same model does a review.