5 ms·
Wow the sentiment here is so negative. I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter. There is
by nl 1mo ago
Wow the sentiment here is so negative.
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
- OrangeDelonge 1mo agoDid it take 18 hours because Fable is comically slow?
- nl 1mo agoI think it's actually Opus 5 that is comically slow! And yes probably that was a factor. It was a lot of code too though.
- yellow_lead 1mo ago> $200 plan (work pays) so violating the TOS? Or work pays for a plan you cannot use at work?
- ec109685 1mo agoWhat part of the tos is violated?
- jsisto 1mo agoyou sound like fun at parties
- nl 1mo agoI have no idea what you think violates the TOS here? I was doing work at work on a work task using a plan paid for by work.
- reticulates 1mo agoThe $200 max plan is for individuals. The individual plans are heavily subsidized. Employers should be using either the Team plan (which has much lower limits than max) or the Enterprise plan (which is entirely billed on usage). Anthropic know that lots of people are doing all sorts of “bad” things like employers paying for Individual plans, (and using multiple accounts to get more usage) and aren’t yet enforcing the rules… but by the letter of the Anthropic terms, your employer should be paying Anthropic a whole lot more (and that’s one of the reasons why AI usage is going to get very very expensive as soon as the subsidies stop, you and a lot of other people are already paying a lot less than you should)
- nl 1mo agoAFAIK there is nothing in the ToS that forbids an employee paying for the 20x, $200/month plan. I could be wrong about this in which case it'd be useful to have a link the clause. I think multiple plans are against the ToS, but I'm not doing that. The Teams plans are more convenient for a number of reasons, but yes, they top out at the 6x plan, not the 20x plan. Edit: ToS are here https://www.anthropic.com/legal/consumer-terms https://www.anthropic.com/legal/consumer-terms and https://www.anthropic.com/legal/commercial-terms https://www.anthropic.com/legal/commercial-terms I've re-read it and I'm pretty sure there is nothing that forbids a business paying for a 20x account. Notably they say this in the consumer ToS: > If you use an email address owned by your employer or another organization, your Account may be linked to the organization's Anthropic enterprise account, and the organization’s administrator may be able to monitor and control the Account, including having access to Materials (defined below). We will provide notice to you before linking your Account to an organization's enterprise account. which goes at least moderately close to indicating using it in a work environment is allowed.
- andrekandre 1mo ago
- pprotas 1mo agoWhat are the consequences? This likely doesn’t matter to most
- epsteingpt 1mo agothe sentiment is negative and justified. anthropic nanny states what you can do. in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore. fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack** but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off. when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out. the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver. business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
- theshrike79 1mo agoSo which frontier AI company still has your goodwill and hasn't lost it?
- aroman 1mo agoI think you need to spend more time with Sol. If you think there is nothing even close to as good as Fable - my guess is you haven’t spent as much time getting as familiar with working with those models as you have with Claude’s. Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents. I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
- gr_norm 1mo agoAgree, Sol and Fable consistently trade blows on Rust dev. I prefer Sol since it's faster, as far as proprietary models go.
- ttul 1mo agoMy sense is that Fable 5 has “taste”. But Sol gets to work and gets shit done. I reserve Fable for when things need a refresh or if I want a flawless front end. Sol does the majority of actual work. I max out two of each at the Max/Pro level every week.
- nl 1mo ago> Codex is more token efficient This is very true, especially vs Opus 5. > tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents. The strength here was less the code quality and rather the long horizon task tracking. This was a very large task - I was chatting with the maintainers and we estimated 4 to 6 months work over multiple phases for human coding. Fable is able to handle that long goal, with incremental steps along the way, handle the verification and course correct when it finds a problem. I think the larger context helps here some, but the strength of the model on this specific thing is notably better. I'm not alone in noticing this. https://www.primeintellect.ai/research/nanogpt-speedrun https://www.primeintellect.ai/research/nanogpt-speedrun shows Fable is able to manage a run nearly 1/3 longer than Sol (8.7 days vs 6.1 days). In my use cases Sol is much closer to Opus 5 though.
- 1mo ago
- nevertoolate 1mo agoI like this story. How did you verify the output? How big is the codebase? Why it took 18 hours? Could you implement it with a small local agent and breaking the task down yourself in two days (i know it sounds like a loaded question, it is not). I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
- nl 1mo agoIt's a differential privacy framework. It's a fairly large code base split across 3 repos. The good thing was that it is fairly easy to verify: we have a working (but slow) version that uses Spark, with lots of existing unit tests. We verified by using those unit tests as well as running our end-to-end process in the Spark and Pandas version and verifying the two databases were within the differential-privacy noise bands of each other.
- jsisto 1mo agohours working on a problem is a weird metric. you could use a slower model and get those numbers way higher.
- nl 1mo agoFair. I was chatting to the maintainers on 2 of the 3 repos this affected and we estimated 4-6 person months work of we were hand coding. The PRs on those 2 repos are 48,000 lines of code (which is a problem in itself!)
- dakolli 1mo agoYou could just write decent specs, or generate decent specs and get this done in 1/5th of the time with smaller models. Complete waste of electricity to run $500k in GPUs full throttle, if not more, for 18 hours straight to migrate from spark to pandas. Maybe try using your brain.
- nl 1mo agoYes, and 4 months ago I'd have done it this way. I'd estimate that generating the specs would have been maybe 2 weeks work? It's across 3 repos, and I'm only really familiar with one. > done in 1/5th of the time with smaller models The coding itself might have been faster, but the end to end time would have been much longer. > Maybe try using your brain. Believe me, my brain was a load bearing seam in this task,
- petesergeant 1mo ago> There is nothing as good as Fable, not even close. This is true, but only for certain tasks. Even as a Fable fanboi, Sol is much better at Fable for some non-programming tasks: Fable for life-planning tasks is miserable because it keeps adjudicating rules, where I've found Sol to be insightful and warm (characteristics I'd previously associated with Anthropic models). We're probably less than a year away from all the frontier models being so good at everything for day-to-day use that it doesn't really matter which you use, which is going to seriously fuck up the business models of all of these companies except the infra companies.
- nl 1mo agoYes absolutely. I should have pointed out that I always prefer Sol for OpenSCAD for example.
- deleted 1mo ago[deleted]
- elAhmo 1mo agoLike others have suggested, you should give more time to GPT models. I sometimes launch Fable with elaborate review personas, it might take 30 minutes or an hour, exceed limits, to produce a review of a PR. Then I ask the same thing GPT without any elaborate 'come up with personas, review the reviews, do rebuttals, etc', and it can find problems that hours of Fable couldn't.
- nullbio 1mo agoEither my $200 sub was getting nerfed, or you're dead wrong about nothing being as good as Fable. The only thing I found it was better at was UI design. The rest, Sol was the clear winner. Refusals, failure to follow instructions, doing 1/10th of the work and then claiming it was "finished" was my experience with Fable. For everything else, there's K3.
- benjiro29 1mo agoI think it really depends on what people expect from the models and how they "code". Sol in my eyes is powerful, but it over engineers so much, that its actually a liability. Where as Opus 5 is slightly under develops but you then can give it a small push for what is missing. I rather have it under develop and i as the human in the loop, can correct/enhance it. Vs the models that adds so much, to the point that your going "dude, stop!". Remember, removing code for a LLM is way, WAY more difficult then adding it. My main issue is with Sol is that its designed to over engineer without thinking why its doing something. Great that you security harden 1000s of lines of code, but ... nobody will ever get to that code. Its that lack of intelligence is where the model becomes a issue for me. Its easier to have less code and then do security audits, with you approving what needs to be changed/hardend. It may simply depend on the developers their mindset. Some folks just want the models to do everything for them, and performance or code bloat means nothing to them (forgetting that this bloat over time makes future LLM work more expensive). Its funny how everybody has their own opinion for what model is better, when in reality its more about that model fits your own development style better.
- nullbio 1mo agoPersonally I'd rather it over-engineer than under-engineer and leave gaps in the implementation that I'm unaware of. Not only does it make life easier to work with the model this way, but over-engineering can be fixed later, as newer models are released, they will get better at cutting out the slop and refining the codebase. The only over-engineering I've really noticed is things like developing extra safeguards and extra tests, which is annoying but not a massive deal. Claude forgetting to implement edge cases I've specifically told it to cover is a big problem though.
- antirez 1mo agoSol for low level programming is consistently better, can work alone for more time, and is faster. If you think Fable is so superior, you need to work with Sol ways more.
- acchow 1mo ago> 2 billion tokens (mix of Opus and Fable), 24 hours That's 1.3M tokens per minute. I suppose you mean input tokens? This would not have cost $2000 at API prices because you would have cache hits.