3 ms·
This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt. I’ve now replaced my use of Opus 4.8
by sothatsit 2mo ago
This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt.
I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
- fastball 2mo agoMedium vs High? Why? From all the charts I've seen the performance jump is pretty large from med -> high (not as noticeable from high -> xhigh).
- sothatsit 2mo agoIf I need something smarter I use Fable. Medium works well and is quick. Opus 5 medium feels much better to me than Opus 4.8 medium.
- dhorthy 2mo agoyeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap
- conception 2mo agohttps://cognition.com/frontiercode https://cognition.com/frontiercode Quality vs cost - medium is the sweet (perhaps better too!) spot.
- ValentineC 2mo agoMedium or low supposedly prevents Opus 5 from overthinking: https://xcancel.com/danshipper/status/2080700057892815114 https://xcancel.com/danshipper/status/2080700057892815114
- rubicon33 2mo agoI think they neutered Fable. When it first came out it was indeed revolutionary. But what we have today, is not what we had before the ban.
- swader999 2mo agoNoticed that too. I wonder if these things just degrade over time, perhaps with the way it writes memories about my project as it goes
- Espressosaurus 2mo agoI’ve observed the degradation, but I suspect what’s happening is they’re tuning it for lower inference costs. Maybe turning down the amount of thinking, maybe quantizing, maybe something else. It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.
- conception 2mo agohttps://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/ their tracker generally shows that isn’t the case. The only times I’ve seen it drop is something broken and just before fable launched.
- nerdsniper 2mo agoI mean they could just be routing known benchmark questions (which all of SWEBench are) to a full-performance variant.
- svnt 2mo agoIs this using the api or using a subscription, though? The incentives are different for each, and it isn't the least bit unexpected that they would maintain API access quality while 'optimizing' the subscription experience to improve their margins (or losses) It seems to do really this you would need to crowdsource it -- users individually give the lab access to a body of subscriptions normally used by average people, and the lab occasionally runs some masked version of the task through on diverse accounts.
- deleted 2mo ago[deleted]
- Balinares 2mo agoCan you elaborate on what felt revolutionary to you about Fable?
- dahdum 2mo agoNot OP, but to me it initially felt extremely proactive and energetic, just powering through roadblocks with ingenuity and enthusiasm. After it came back I was constantly getting refusals and downgrades for things Opus had been doing. I’ve written it off for my use cases and getting by just fine with Opus 4.8 and now 5.
- sothatsit 2mo agoI got Fable to run overnight and I woke up to a working prototype of a very complex feature. And then I did it again for another complex feature the next night. The code still took weeks to clean up, but it worked and was correct. It felt then, and still feels, like a big step change on very hard problems. These are problems I would previously expect to take a month or longer to implement. I have also noticed Fable can handle much more nuance when reasoning through writing and research, but that is harder to quantify.
- slopinthebag 2mo agoWhat was the very complex feature?