4 ms·
I couldn't find a situation where Fable was significantly outperforming Opus enough to make me say wow. I tried to make it fix a browser game that is sort of l
by _pdp_ 3mo ago
I couldn't find a situation where Fable was significantly outperforming Opus enough to make me say wow.
I tried to make it fix a browser game that is sort of like a Mario clone. It couldn't. It fumbled in the same places Opus was struggling too. I tried it with other code as well, but I couldn't get any significant performance improvement out of it, except perhaps in improving my account's token burn.
If anything, in my opinion, GLM 5.2 had a better moment than Fable recently. Not because it is better, but because it was not hyped at all, and many people realised that it is possible to run a serious open-weight model yourself, as long as you can get the hardware to support it.
I am not drawing a direct comparison here, because Fable is clearly the better LLM. But GLM 5.2 is a good, honest model, and I think open-weight models will only get better going forward.
GPT 5.6 is claimed to be at a similar level, or even better than Fable. We will see. They don't seem to hype it as much, and I have not read anywhere that anyone found a soul or consciousness inside it. And if it benchmarks well, I would possibly use it more for this very reason.
It reminds me of that story from Nassim Taleb's Incerto series where if you have two surgeons at practically the same level, but one looks like the typical surgeon and the other looks like a butcher, who are you going to choose? Taleb suggests that the answer should probably be the butcher, because to get to the same level while looking the part so much less, they probably had to be much better than the data shows for.
I cannot also understand the hype online claiming that the Fable transitioning to token-based billing after the gratis period is equivalent of being in the permanent underclass. The only impressive demo that I saw was it writing NES games which kind of looked fun but I couldn't find more details and I am not sure if you can get this done with another model - probably you can but nobody is trying.
So great model but it does not have the same effect as Opus 4.5 and Codex which made me feel that there was a stepping-stone change.
GLM 5.2 had that moment though.
- scotty79 3mo ago> many people realised that it is possible to run a serious open-weight model yourself, as long as you can get the hardware to support it That's not realistic. You'd need non-consumer hardware for a frontier open-weights model. And even if you had such hardware for free. Electricity to run it would cost you more than a sub.
- _pdp_ 3mo agoWhen I said "running it yourself" I didn't mean me personally running it at home. That will be unfeasible. A company can afford it though and also a supplier that serves multiple customers can do that as well. In fact, I was talking to a company that does this and the model is fully in UK where we need them to be. At the end of the day it is about the sense of optionality. We know that this is the business model for almost all open source projects. It is not like you cannot download and run the project yourself and some do for practical reasons, but often times the cloud version is priced such that it is the path of least resistance so people go for that.
- randsorex 3mo agoI agreed last week but the more I have used Fable this week the more I am getting use to Fable. I think the Taleb argument is really a stretch for this situation. I consider Taleb one of my greatest teachers but that kind of Talebism I have grown suspicious of in time. How many surgeons actually look like a butcher lol. Much of Taleb is like this that it sounds profound as a thought experiment but so much has just nothing to do with reality. A lot of what he is arguing against he is actually doing a form of. I love Taleb but he is absolutely full of shit. It is marketing and his brand. The proof he is correct are some vague options trades he made 40 years ago.
- _pdp_ 3mo agoFable is better than Opus and I do not think there is much argument there. And yes, it is natural to get used to a new tool and even end up liking it more over time. What I am trying to say is that it does not feel like the step change they are claiming it to be. I fully agree with your point about Taleb. And yes, it can come across as a bit glib, but the underlying point is not new. Do not judge a book by its cover. That idea exists across many cultures, fables, and stories, so it holds IMHO. The surgeon example is just an illustration of that idea, and it is a great example because it makes the point memorable. In practice, of course, I agree that most modern surgeons are trained to broadly similar standards and look and act the same, with some outliers and historical exceptions.