4 ms·
Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.
by johnnyApplePRNG 2mo ago
Fable is most definitely nowhere near 10T.
The cost to train and infer that would be insane, even by today's standards.
- nl 2mo agoFable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai-model-nearing-anthropics-mythos-ft-reports-2026-08-07/ https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.
- andai 2mo agoWasn't Opus ~1.5T and Fable is about twice that?
- stymaar 2mo agoThe fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac. Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.
- WinstonSmith84 2mo agoYes. And Opus goes a very long way compared to Fable, Anthropic isn't doing any favour, it's clearly just 2 models with a very different amount of parameters.
- nl 2mo agoThe open models don't really match Opus. For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention. I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.
- stymaar 2mo agoEven if they don't match current-day Opus in everything, they do beat 6 month old Opus, which we have no reason to believe it was smaller than the latest version.
- johnnyApplePRNG 2mo agoIf Fable is seriously around 10T and Kimi K3 sidles up to it at 2.4T That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.
- nl 2mo agoIn real world comparisons K3 is somewhere between Sonnet and Opus. Fable is just a completely different (higher) level. The long tail of tasks and queries is where you see the difference.
- riknos314 2mo agoKimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.
- zozbot234 2mo agoThe raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you're paying for the model includes that raw margin.
- alightsoul 2mo agoThe Chinese have similarly high profit margins on Inference via their first party api
- x-complexity 2mo ago> The cost to train and infer that would be insane, even by today's standards. This assumption is likely what has led to the erroneous failure. Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains & datacenter scale increases, even 50T+ is well within reach at the top end.
- kube-system 2mo agoYeah, that's insane, you'd need to have many billions of dollars and buy up a huge chunk of the worlds memory supply to do that /s