8 ms·
This matches my experience. Burned $2K to see how it will perform on frontend tasks and backend tasks. Frontend did a significantly better job than Opus on toy
by renoir 4mo ago
This matches my experience. Burned $2K to see how it will perform on frontend tasks and backend tasks.
Frontend did a significantly better job than Opus on toy-scale wireframe projects by using gimmicks like fluid dynamics. Then when given medium to big tasks like multi-page web app where layouts and aesthetics must be decided by model itself, results by Fable and Opus scored indistinguishable score from human judges.
Backend, gave tasks related to setting up a data flow that involves Postgres, R2, Kubernetes, gVisor, so on. The noticeable gap was, Opus did better than Sonnet, but Fable actually returned a result that fails and confidently stated it ran X, Y, Z tests to ensure it works and got these results. Very surprising, given neither Opus nor Sonnet suffered such problem.
Longest frontend task was ~2H. Backend, 8H.
Though none of the tasks were related to developing LLMs, (just production grade secure system that could've been developed 20 years ago, no LLMs involved), it is possible Claude Fable downgraded itself or spitted out fake results. There'd be no way of knowing since Anthropic silently degrades model quality based on undisclosed internal criteria which claims to be about LLMs.
We decided Fable is unpredictable and cannot be trusted to the degree that Opus and Sonnet can be trusted for any projects beyond toy-scale quick wireframes, but Fable can be the best tool for quick UI UX wireframing for non-technical roles.
- weatherlight 4mo agoI had almost the opposite experience. I'm building a compiler for a language without a tracing GC, so a big chunk of the work is around memory management: functional in-place update, reuse analysis, and a Perceus-style reference-counting strategy similar to what Koka uses. The hard part was that my use case wasn't exactly covered by the Koka/Perceus paper. The prior art got me maybe 75% of the way there, but the remaining 25% was a cluster of bugs with very similar shapes and no obvious published solution. With Opus, I kept getting stuck in this loop where it would fix one case, but break another case elsewhere in codegen. We ended up with something like 16 failed experiments just for one bug class. The workflow was: run an experiment, identify the shape of the bug, propose a fix, check whether it emitted the correct Zig, then see if the fix broke any previous memory-management cases. It was useful, but it kept choking on the parts where there wasn't clean prior art to lean on. Fable was a different story for me. It one-shotted the Class A bug cluster, and then basically said "by the way, your previous attempts have these structural problems." More importantly, it identified the other related bug classes and came up with workable strategies for applying the Perceus-style memory management in those shapes too. That's obviously anecdotal, and I'm not claiming Fable is universally better. But in my case, this was not a toy frontend wireframe. It was compiler work involving ownership, reuse, RC/drop behavior, and Zig codegen. The thing that surprised me was that Fable seemed better precisely where the problem wasn't just "reproduce known prior art", but required filling in a missing piece. Also worth noting: I'm not using the API. I'm using the Max plan, so maybe there are product-path differences here. But I definitely did not have the "unpredictable beyond toy-scale" experience. For this particular compiler/memory-management problem, it probably saved me a ridiculous amount of time and money.
- cmenge 4mo agoSimilar. I gave it a really hard task, basically messy code in a complex domain that was bug-ridden from a mess previously created half manually and half by Opus. It cleaned things up beautifully, both the backend and the frontend. Maybe the prompt was particularly well-suited for the model (I instructed it to put on a mathematician's hat, look at the mathematical substructure of the problem, identify invariants and general laws and verify them, then plan how to remediate). It wrote a ca. 800 line in-depth analysis (at times spawning over 130 research agents...) with remediation plans, prioritized them and then implemented them. One issue was that this document was frankly over my head. Both the language it used and the mathematical parts were very terse, and in parts it felt like a post-C2-vocab exercise. The prose was much harder to understand than the code snippets / data models. As a non-native speaker, it lost me on the prose part, and had to ask it for a less elaborate version to actually understand it. It burned the session limit four times, but it turned a huge mess of proof-of-concepts with patchy glueing into a coherent, stable application. I'm also on the Max plan using Claude Code, and I have the feeling that the harness is much more important than the consensus expectation.
- ElFitz 4mo ago> and I have the feeling that the harness is much more important than the consensus expectation. Is that really the consensus? There’s been a bit of literature lately on that. Can’t find the one about looking into whether or not the harness had a greater impact than the models (for comparable models), but there’s this one: https://arxiv.org/html/2605.23950 https://arxiv.org/html/2605.23950
- selimthegrim 4mo agowhoa, my university!
- comboy 4mo ago'by the way, your previous attempts have these structural problems." Just to be clear, it did not have access to any previous work that opus did? Because they are pretty good at digging out relevant tmp files and making use of whatever is out there. With my fable adventures I caught it hallucinating something and stating it as a fact in CLI twice. And it was something that I did not see opus do in such way, opus obviously many times stated some things that it did not verify but guessed, but fable said something like "the probe showed that ..." - but there was no probe, it was not about some past events it was about what it was doing right now. "I overstated"... But boy does it know Chinese, so much better than any other english model, gemini used to be the king but fable clearly was trained on a decent amount of it. It has a deep cultural understanding.
- deleted 4mo ago[deleted]
- alasano 4mo agoThere's an often hard to express subjective experience you get with a new model, especially if you spend a lot of time trying out different ones. I believe the people who feel like Fable is a big improvement, for me it's just much more reasonable and grounded. It makes me realize how much of a try hard over optimizing planner GPT 5.5 can be. I've been fighting it often to simplify plans. But no matter the model you can't trust them to actually deliver on very long tasks while maintaining quality. At least not without external orchestration and review.
- espeed 4mo agoRun /model after your task to see. Mine keeps downgrading to Opus 4.8, which is a problem because Opus 4.8 keeps no-oping critical security code.
- comboy 4mo agoThere is in /config "Switch models when a message is flagged" now which can be set to false, but I had no chance to see what happens then, does it just stop or what.
- deleted 4mo ago[deleted]
- espeed 4mo agoSession paused Fable 5 has safety measures that flag messages on most cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Send feedback with /feedback or learn more 1. Switch to Opus 4.8 2. Edit prompt and retry with Fable 5
- staticautomatic 4mo agoBiology? Why?
- adgjlsfhk1 4mo agothey're worried about people creating bioweapons
- tekacs 4mo agoWhat you're describing only applies to security or biotech downgrades. A downgrade related to the model believing that you're doing something related to model development is invisible and silent and internal.
- jasondigitized 4mo agoA single 8h task? I'm sorry, but that's just asking for trouble.
- queuebert 4mo agoI don't understand how some of y'all use these things. I get garbage unless I give them very specific concrete tasks with as much context as possible. Anything that takes more than 30 min is usually a waste because the scope was too large.
- standardUser 4mo agoYou have to build up a context, or otherwise seed the memory, to get anything useful out of these LLMs on a large or existing project.
- whstl 4mo agoDifferent people just have different concepts of what's garbage and what's not. There seems to be some kind of AI hysteria going on, with people becoming so enamoured with the AI that they accept anything it produces as if it's some gift from the gods, while others just reject it prima-facie. For example, the worst design I have seen recently was from a designer who pivoted into "vibe coding influencer". The worst code is from developers who were heavily into Clean Code a couple years ago and now half their PRs is unused dead code.
- gessha 4mo ago“One man’s trash is another man’s treasure.” takes a new meaning in today’s agentic coding world.
- smoe 4mo agoI had good experiences doing multi-hour refactoring/housekeeping tasks that basically consisted of applying the same steps and rules n times. Worth noting, a significant chunk of those runs involved the agent waiting for the compiler, linters, type checks, and test suites, as well as updating journals. It’s not the agent sputtering out code for eight hours straight. And naturally I spend more time on manual verification in the end as much less of it is happening during the coding process.
- standardUser 4mo agoAt a certain point, people value reliability over improved performance. I think a lot of us have hit that point as this technology becomes indispensable to our work. I'm sure I'll use Fable... eventually. But at 2x the cost, I'll skip the inevitable learning curve for now. And thanks for your insights! Not surprising to me that any new model would, as this juncture, be more cryptic and inconsistent than the current models.
- hirsto 4mo agoThis seems insane to me. Aren't long running tasks an anti pattern at the moment? My understanding of literature is that small mistakes in chat history cause a trend away from performance
- colechristensen 4mo ago>Aren't long running tasks an anti pattern at the moment? Longer running tasks require better setups and several ways of pinning the progress to reality. When you have that though things are quite all right. A good long running task will run inside a framework that it's not trying to modify.
- dwaltrip 4mo agoFable is a lot like Opus at its best. It's simply more reliable and feels a bit smarter. For my use cases, using it feels very nice, and notably better than Opus. It needs less direct guidance to get reasonable looking code and I don't have to watch it as closely. For context, my Claude Code working style is quite heavy on discussion "to align" before implementing anything. We also use a good amount of Markdowns. Oh yeah, it also is has way less "phrasing quirks" and is a clearer communicator. Opus 4.8 was a bit of loon with some of its writing styles. I had mostly straightened it out, but not entirely. It would use the most ridiculous flair at times.
- winrid 4mo agoI've had Fable add Chinese characters to our conversation for no reason.
- taikon 4mo agoSame here
- maaaaattttt 4mo agoCould it be that Anthropic is using the Chinese characters trick to consume less tokens behind the scenes?
- elbear 4mo agoI've only had that happen with Chinese models until now. Interesting that Fable is doing it too.
- 4mo ago
- skerit 4mo ago> Burned $2K to see how it will perform on frontend tasks and backend tasks Burned $2K on some kind of enterprise account or ... ? Why not just get a $200 Max Pro account? While I'm loving the output of Fable 5, I will *never* pay the "normal" API token price for it. You can reach $2K in a stupidly fast amount of time.
- unholiness 4mo ago> I will never pay the "normal" API token price for it. Not until June 22 you won't!
- KellyCriterion 4mo agoCurious: >Burned $2K In which time was this burned, because it sounds like "I gave it just a bunch of menial tasks to solve" - or did it run for like 1 complete day continuously?
- aleph_minus_one 4mo ago> Burned $2K to see how it will perform on frontend tasks and backend tasks. When I read such statements on HN, I nearly always ask myself: if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs?
- graphime 4mo ago> if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs? Do you think US$2,000 is a lot of money?
- inferiordev 4mo agoIt is lot of money to burn.
- mystifyingpoi 4mo agoIf I pay for it, yes. If my employer pays for it, no.
- bdangubic 4mo agothat is much better spent money by employer than to give you extra compensation. but as you said, not a lot, who needs $2k after all
- nrjames 4mo agoI'll bite. Yes, it's a lot of money. It's several months worth of nice healthy groceries for a family of 4. It's my annual deductible on my health insurance. It's slightly lower than my annual property taxes.
- mym1990 4mo agoNow that we have trillionaires running around, it may not seem like it, but it is a considerable amount of money in most of the USA. In many parts of the world it would be considered an unfathomable amount.
- nullbio 4mo agoI genuinely think that Fable is just Opus 4.8 with some extra skills and harness. I saw a video of someone generating UI with them both side by side, and it gives identical recommendations for themes etc. Doesn't feel like a new model to me, just Opus 4.8 with some sprinkles on top.
- aspenmartin 4mo agoThose are some incredible sprinkles.
- Resoluciones 4mo ago[dead]