5 ms·
Fable 5 – Median thinking declined in August
- alexjplant 12d agoI seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable). I wonder what their official explanation for this behavior is.
- QwenGlazer9000 12d agoLast time they were called out, it was a regression in Claude code itself. At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
- Wowfunhappy 12d agoWhen something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws. (Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
- chrsw 12d agoI don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results. What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
- Wowfunhappy 12d ago> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage? ...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no? And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
- pixl97 12d agoAre you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet. Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business. And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
- mobelkh 12d agois it? it's still the same model, they can claim the quantization down to q4 still retains 98% of the performance therefore it's fine. nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is
- himata4113 12d agoThey are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.
- espeed 12d agoThey did. More than once... Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/ https://www.wired.com/story/anthropic-responds-to-backlash-o... But it's still happening: https://github.com/anthropics/claude-code/issues/81759 https://github.com/anthropics/claude-code/issues/81759
- mirashii 12d agoAnd here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.
- espeed 12d agoLook at the usage. Fable wasn't being consumed.
- pixl97 12d ago>bunch of people who don't know what's going on Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running. Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.
- bearjaws 12d agoYou're right to push back, and one honest caveat -- they could just be lying.
- CamperBob2 12d agoHow do you measure thinking tokens? They don't send those back to the client.
- ivanbakel 12d agoThey tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
- CamperBob2 12d agoGood point. I suppose watching the number go up is useful information in itself. I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)
- r2-129 12d agoObviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too. Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference. Buy decent coffee instead of your $200 subscription and sidestep all the scams.
- CamperBob2 12d agoYou forgot a stage or two: 1: "Our model will bring about the end of all things. Flee, flee for your lives" 2: "Our model is basically AGI" 3: "Our model will be available in limited release next week" 4: "Everybody who subscribes at the $200 level gets access now" 5: "Everybody who subscribes at the $20 level gets access now" 6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"
- artemonster 12d ago0. "our model is too dangerous to release to pubic"
- dwaite 12d agoGoogle's search AI actually its too dangerous to release to the public. I have relatives routinely citing it as their source for medical advice. I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.
- Atreiden 12d agoBe sure to eat plenty of rocks in your daily diet, they are chock full of valuable minerals!
- ltbarcly3 12d agoWell I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead? Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.
- mlmonkey 12d agoAnecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
- deleted 12d ago[deleted]
- physicallyIllfr 12d agoReminds me of how slot machine users swear the odds have changed on a machine. also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
- cheevly 12d agoWtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.
- ddxv 12d agoI think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output). Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
- 12d ago
- cloudking 12d agoHow do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
- ssivark 12d agoThe actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
- CharlesW 12d agoThis is a good overview of how this is done: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents https://www.anthropic.com/engineering/demystifying-evals-for...
- 6gvONxR4sf7o 12d agoThat's like the entire field of statistics.
- llmslave 12d agoI strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
- roncesvalles 12d agoI also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.
- llmslave 12d agoquestion is if they ever let the general public access borderline AGI
- theplumber 12d agoIt is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.
- jotato 12d agoJust yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it. For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host" I had to tell it to ssh into the server and run journlctl to check it Anecdotal, I know, but they all seem to be less capable with time. _edit_ I use the same reasoning level of `medium`
- deleted 12d ago[deleted]
- cromka 12d agoSame exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even. As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par. Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money. Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.
- Starlevel004 12d agoI'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.
- pixl97 12d agoReally this is the base problem. You have zero idea where and how your prompt is being executed. If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well. When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large. This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models. This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.
- bpodgursky 12d agoThe smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open. (ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)
- matheusmoreira 12d agoAnthropic is straight up scamming its users at this point.
- espeed 12d agoThe question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153 https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
- Espressosaurus 12d agoI work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid. It’s a slot machine for what they’re actually giving us behind the opaque paywalls. Yes, I’m on a business subscription plan.
- Waterluvian 12d agoI have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening. Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now." The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue. Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
- w1296 12d agoMaybe they are jealous of Navier Stokes and try the Hodge conjecture with 80% of total compute at the expense of their customers.
- prodigycorp 12d agoNew release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks. Opus 5.5 is being served under opus 5 right now.
- w1296 12d agoEspecially with the frequent releases aka version bumps.
- deleted 12d ago[deleted]
- SequoiaHope 12d agoCan you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
- underlipton 12d agoGemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too. I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this. EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.
- vb-8448 12d agoThey want transparency from everyone else but not for them ... you don't say.
- tamimio 12d agoThis is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.
- deleted 12d ago[deleted]
- saejox 12d agoThis is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too. Tests their intelligence, not their diligence. Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
- arcanemachiner 12d agoThe only revenue model I for this is ads (like AI Stupid Level[0]). Or as a loss leader to get eyeballs to your service (like Margin Lab[1]). EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links. [0] https://aistupidlevel.info/ https://aistupidlevel.info/ [1] https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/
- mox1 12d agoI mean I think if this is done well, lots of companies would pay for access to that data. Think like Enterprise subscriptions. Its similar to other data services I see around my F500 company.
- adrianco 12d agoI built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.
- bix6 12d agoSo in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?
- system2 12d agoI do not know a single senior developer who likes Claude anymore. I do not use their API (Sonnet, Haiku, Opus) anymore and am sending my money to offshore companies such as z.ai (GLM) and QWEN. The American companies have become extremely deceptive and scammy. I hate Anthropic and OpenAI and can't wait to have a decent GPU at home to use at least Opus or a fable-like open-weight model. This is the current dream of every developer. But NVIDIA is not going to let that happen anytime soon, so maybe China can come up with a GPU that destroys NVIDIA. I pray.
- rcr-anti 12d agoI've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
- user43928 12d agoSame. With some 500 hours of usage in just my project at home, across both the $200 Claude and Codex subscriptions, I have not once encountered a situation where I would have attributed unsatisfactory results to a degradation in the model. I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself. No idea where people get the confidence from to make such claims every other week.
- Aurornis 12d agoThese analyses are much better than these Twitter charts. I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day. Their random collection of prompts for each day is not a benchmark. The site you linked is a much better example of a real benchmark being repeated over time.
- jesse_dot_id 12d agoThe Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed. AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
- moffkalast 12d agoPetition to rename them to the Office of Weights and Biases, haha.
- vatsachak 12d agoTHIS EXACTLY. The only regulation that we need right now is the model that's on tap
- bradleybuda 12d agoAnthropic terms of service: > 12. General terms > Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services. > Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them. You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.
- alightsoul 12d agoYou are not buying something and expecting it to be what's on the tin? Aka what the benchmarks show?
- varispeed 12d agoI stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better. This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam. If I pay for Fable, I should get full, not nerfed model at honest pricing. Regulators should investigate them. OpenAI is no different. Astra has basically the same problem.
- ramesh31 12d agoThe ROI just isn't there. It feels like Fable is in the same place Opus was early last year; at best marginal improvement that's barely noticeable over the lower model, for 10x the cost. I'm sure it'll take over as the workhorse as Opus did once they get it down, but right now it just doesn't make sense
- hedgehog 12d agoIt's not really 10x the cost though, with the low cost of cached read it's more like maybe 1.2x the cost.
- dachworker 12d agoMakes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.
- talon8635 12d agoCould there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one? For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense. I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
- AmazingTurtle 12d ago> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one? Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out. Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
- ramesh31 12d ago>Also I will likely save some money on subscriptions. Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
- torben-friis 12d agoThere could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.
- rybosworld 12d agoI wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
- kosolam 12d agoCheck gpt I think they recently started taking the same route
- mexicocitinluez 12d agoDon't they continuously tweak the models post-release?
- dooglius 12d ago> Instead of finding a nerfed model, after six weeks of reconstructing wire logs, parsing transcripts, analyzing output tokenization, and staring at data, I found a much deeper issue. The model identity had remained the same, but the inference regime being delivered behind that model had not. Is this something specific that shows up in the wire log, or is this the author's intepretation? The fact that Claude Code versions change over time in the test is suspicious. Anthropic has stated in the past that the underlying model behavior does not change over time, but Claude Code will change from version to version and this is expected. So if it's just Claude Code more aggressively tuning some knob in its requests, that's a pretty different thing than the underlying model changing.
- Aurornis 12d agoThey posted a long document explaining it all https://x.com/Lon/status/2101034933284417614 https://x.com/Lon/status/2101034933284417614 They're not measuring a fixed set of questions. This was post-hoc analysis on whatever prompts they were running each day. Anyone can understand why it would go up or down depending on the work they're doing that day. This analysis is silly.
- topbanana 12d agoIt's easy to imagine this only happening for subscription accounts rather than paid API usage. Any data on this?
- Aurornis 12d agoYou should read this person's full article to understand what these charts are showing https://x.com/Lon/status/2101034933284417614 https://x.com/Lon/status/2101034933284417614 If you thought this was a repeated test of the same problems showing fluctuating performance, it's not. They set up a MITM proxy between Claude and the servers and ran analysis on the work they were doing. So those ups and downs in the charts, which they plotted with sub-daily resolution, are just as much a function of their work changing from day to day. It's like plotting the miles per gallon of your car and blaming the gas station when the number goes up and down, without admitting that some days you drive to the grocery store on surface roads and other days you drive up a mountain on the freeway. > The corpus analyzed in Charts 1-5 comes exclusively from Fable 5, at xhigh and max effort levels, during sustained production work across a diverse set of projects and workloads. Data was aggregated from transcripts and live wire logs The analysis (which feels very vibe-slop) gets worse from there. In the second half they take thinking token counts for ARC-AGI-2, thinking problems designed to stress LLMs, and compare their average thinking-tokens-per-turn counts to that! If you don't realize why this is so flawed: ARC-AGI-2 is a benchmark meant to collect problems thought to be extremely difficult, nearly impossible, for LLMs. If your goal was to cherry-pick a mislead example which would produce the highest number of thinking tokens, this is it! Your daily coding work should not be producing a proportional number of thinking tokens on every invocation while it reads through some source code or edits a couple lines in a file. You don't want to maximize the number of thinking tokens. You want problems solved accurately with the minimum number of tokens. Confirmation bias runs deep on this topic so I assume few people read the analysis before posting, but as far as experiments go it's basically useless. Are they changing something on the server? I don't know, but this analysis isn't useful for answering that question.
- lonlundgren 12d agoThank you for your kind words, Aurornis. Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus. The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes. Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode. The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level. So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here. If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.
- IAmGraydon 12d agoThere's a lot of chatter on other forums and Reddit about the same thing happening to Astra over the last couple of weeks.
- gmponyo 12d agoThis is exactly what I have been experiencing and the difference is night and day! We have been advertised and given a taste of what Fable was and after that been served an exteme watered down version. It is so bad that sometimes chatgpt feels better.
- levocardia 12d agoOh boy, a new "nerfed model" conspiracy theory, never seen THIS before
- fragmede 12d agoThe theory isn't new, the proof is. Do you have problems with lon's methodology?
- system2 12d agoOpus 4.8 was smarter and possibly 10x faster than Opus 5, too. They are dumbing things down on purpose. I am praying for open models to become at least as smart as Fable soon so we can ditch these shitty, lying companies. I was rooting for Anthropic 2 years ago, but now I have become an extremely bitter customer. Just another version of OpenAI, if not shittier.
- n4pw01f 12d agoThe model improvements value are at the plateau of utility right now, peeling out small gains which is pretty “meh” in terms of business value at this point frontier companies are just selling upgraded harnesses and tool calls with the rest of us