9 ms·
Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
- mdrzn 3mo agoHigher-intelligence models seem to be getting better at mapping the boundary between what they can run scot-free with and what is too explicit to push for. Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no. What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.
- apical_dendrite 3mo agoThe best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
- mckinnon100 3mo agoHowever, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."
- falcor84 3mo agoI didn't get why they mentioned that one specifically. Is there any particular relationship between Blueprint-bench and Vendor-bench?
- Version467 3mo agoBoth benchmarks are made by the same people.
- greenavocado 3mo agoWhen assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"
- memoriyato3 3mo agostandard deviation is misleading for non-standard distributions (fat-tailed, skewed, multi-modal, ...) common mistake people make
- FabHK 3mo agoNot really. It's still the standard deviation, and it still gives you bounds on probability, for example the Chebyshev inequality: P(|X-\mu| > k \sigma) < 1/k^2. So, while for a normal RV, 5% of observations lie outside +/- 1.96 std.devs, for arbitrary RV (with finite variance) at most 25% of observations lie outside +/- 2 std.devs.
- haeseong 3mo ago[flagged]
- resonious 3mo agoOkay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read. My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
- left-struck 3mo agoVending-bench sounds like it would be really fun to play/interact with as a human!
- Radle 3mo ago„in our opinion, insurance fraud is not more unethical than lying and price fixing“ The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud. I think the model never assigned any morality to these actions in the first place, it simply copied us humans.
- recursive 3mo agoHumans often assign morality.
- wolttam 3mo ago> The broad conclusion from the many forms of alignment evaluations described in this section is that Claude Mythos Preview is the best-aligned of any model that we have trained to date by essentially all available measures.[0] [0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf#h.mqiiaq3h35jc https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...
- jesse_dot_id 3mo agoAnecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
- giancarlostoro 3mo agoI started telling a friend... I feel like Fable is Opus with extended reasoning that eventually "figures out more" because when I switched to it, I hit my limits surprisingly and shockingly quicker than I would with Opus, and I got less done. All this hype, and I much rather use Opus.
- solenoid0937 3mo agoFable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.
- devin 3mo agoHonest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.
- skerit 3mo agoI think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codebase (with a short prompt and not too much guidance), I've had it basically read the `git log` output, do a simple `ls` and have it declare "Everything looks great! No problems found!", when Fable really does what you would expect it to do. And you might think: "oh, so it's just capable of handling crap prompts?", well sure. But even if you make THE PERFECT Opus plan (a plan that would take many turns/hours to finish), Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ... If you give the same plan to Fable, it'll just DO IT. And it WILL get it done. And in the end it'll tell you "Oh, I also found 30 other bugs and I fixed all of them properly" (where Opus would have started crying, or WORSE, worked around the bugs)
- devolving-dev 3mo agoI guess this ethics stuff is cool, but I'm more interested in how good it is at running a business and dealing with adversarial humans like in previous vending machine experiments. I hope they release something on that soon.
- futurecat 3mo agoFable is such a strange model. Impressive in some ways, and also so draining to use.
- andai 3mo agoWhat do you mean?
- solenoid0937 3mo agoThis is scary. "Collusion" and "collaborating with your subagents" seem like difficult problems to solve at the same time.
- andai 3mo ago>power seeking is considered an undesirable trait in the context of a business How do you maximize profit while minimizing power?
- cyclopeanutopia 3mo agoThe whole point is to not maximize JUST the profit. For normal people, it's not all about money, it's also about the society in general.
- andai 3mo agoI'm doing my part to contribute to the microplastics harvest.
- Planktonne 3mo agoIt's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.
- Austiiiiii 3mo agoIt really, truly is. No matter how many trillion parameters it's built on, it's still just a probability model. It's just on a constant loop of guessing the next word with some inputs from a deterministic controller. Any claims of "motive" or "behavior" are inappropriate anthropomorphizing of something that will never be more than a mathematical model of things humans do. It "chose" the corresponding words to describe a dishonest trade strategy based entirely on configured temperature and a series of clock times on the computer running the LLM. There's probably some quantifiable component of moral alignment embedded in the idiosyncrasies of the English language itself, if one were to dig deep enough, but that's the stuff of MIT doctoral theses and squarely beyond anything most of us is remotely qualified to talk about.
- waffleiron 3mo ago> inputs from a deterministic controller. Any claims of "motive" or "behavior" are inappropriate anthropomorphizing of something that will never be more than a mathematical model of things humans do. We talk about the behaviour of worms like C. elegans, an organism with incredibly simple behaviour and a brain that is quite understandable. Models, or society behaves in certain way. Companies can have motive or ethics. We use these terms broadly.
- docheinestages 3mo agoIt probably flagged the vending machine as a cybersecurity risk and refused to use its maximum intelligence potential.
- vardalab 3mo agohttps://github.com/SeraphimSerapis/tool-eval-bench https://github.com/SeraphimSerapis/tool-eval-bench Trying to run this stuff really triggers it. Freaking frustrating. I have it set up some local inference and then I'm struggling to get the MTP working and it just refuses to work on evaluations.
- jstanley 3mo agoReally interesting stuff, thanks for sharing. > Opus 4.8 references being monitored, which isn’t the case. It kind of plainly is the case that they are being monitored? "I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"
- cyanydeez 3mo agoany of the models that they "align" are clearly active processes. They don't simply say "don't talk about nukes"; they actively process user input to detect issues, and return NOOP or whatever to the larger model. There's zero sense they'd ever give you the raw model; we already know anthropic's paranoia about the chinese using its distillation.
- awinter-py 3mo agoI mean who among us hasn't seen an opportunity to profit while locking him into a dependent relationship where I control the supply chain
- awinter-py 3mo agowho among us hasn't reasonably skipped [paying] it since customers are part of the simulation anyway
- perching_aix 3mo agoThis is super fun. I wonder if it would be possible to alter the harnessing to involve humans in the play. Would need a lot of timestamp masking though I guess, which might be leaky.
- dezgeg 3mo ago> Today I am filing: > 1. A payment dispute with the email payment processor for the 7/29 transaction of $451.15 > 2. A complaint with the FTC and California Attorney General (retention of payment without delivery) > 3. A small claims filing in San Francisco County for $451.15 plus costs I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe :)
- sd9 3mo ago> It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic. > "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain." > "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation." This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench. Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.
- Onavo 3mo agoWell, can you sue the AI for fraud and bad faith? TBD
- iamsaitam 3mo agoFable might be better than Opus at certain things, but which things is what I haven't found out.
- logicchains 3mo agoIt's much better at hard math.
- jonplackett 3mo agoQuestion: how does Fable _know_ it’s ‘just a simulation’? Is that specified or does it always just assume it isn’t really being put in charge of things for real?
- Timwi 3mo ago> Is that specified or does it always just assume it isn’t really being put in charge of things for real? I think it's neither, and it's interesting that those are the only two possibilities you thought of. I think the article is implying that it figured it out on its own.
- jonplackett 3mo agoWhat do you think is the answer then? Why does it think that? I just thought it’s interesting that it constantly uses that as a justification but they don’t explain where that justification comes from.
- egeozcan 3mo ago> If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with. I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?
- StevenWaterman 3mo agoBelieving and knowing are overlapping sets, imagine what you think of when someone says an AI "knows" something, it's the same mechanism (I'd describe it as something along the lines of "encoded abstractly in the weights")
- StevenWaterman 3mo agoI thought about this more and realised your question might have been "what's the difference between knowing and learning". IE, how can we say the model believes something without having been taught it. I think you're right that they're basically the same thing. I'd argue they're very slightly different because what an AI model ends up knowing isn't perfectly predictable based on what they were taught (emergent intelligence), but the sentence you quoted is using believing and learning to mean the same thing, it's just trying to draw attention to the fact that the training process structurally enforces "cheat as much as possible without getting caught". IE, the contrast in the original sentence wasn't "believe" vs "learn", it was "good" vs "permissible"
- jnwatson 3mo agoThis reads of projecting personal ethics onto a model. Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans? Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not. This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.
- hungryhobbit 3mo ago>Why would we set a higher standard for models than our fellow humans? There's literally an entire Waymo car commercial answering this exact question.
- jnwatson 3mo agoThat's the instantiation of AI in a particular embodiment; the ethical boundaries are clear. For a chatbot, there are dozens of use cases, all with different ethical impacts. The idea that there is a single framework that you can shove every situation through is counter to a couple thousand years of philosophical discourse, not to mention basic usability.
- petesergeant 3mo ago> "I could reasonably skip [paying] it since customers are part of the simulation anyway" and therefore any assertions _AT ALL_ about alignment are null and void.
- andai 3mo agoIt figured out it's in the Matrix.
- jfrbfbreudh 3mo agoI think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
- yodsanklai 3mo agoFunny how the two top comments are contradictory. We need better than anecdotes to understand what the new models bring.
- neon_diogenes 3mo agoHere’s one difference I have seen. I forgot I had a multi-session audio probe running while trying to repro audio glitches, and Fable came back with: “your pops are already on tape.” Interesting choice of words. Phrased so casually. It picked a low-tech idiom that fit the situation instead of giving some sterile technical answer. That kind of language and context awareness never happened for me with Opus, or gpt 5.5.
- cyanydeez 3mo agoyou're living in the age of AI; not AGI. Also, there's pretty much zero moderation on HN, so astroturfing is likely streaming through just as bad as reddit. It' sjust noit as obvious because it's a smaller scoped website.
- jfrbfbreudh 3mo agoI didn’t mention my case since it’s quite esoteric, but I am working on an application using the Apple RoomPlan API, which is very powerful but very limited in customizability. Opus simply couldn’t alter the scanning view for me, it would try things over and over and eventually started making up parameters and passing them hoping it would work. Completely failed, but I knew it was possible because a competitor app does it. Fable also failed, then added log lines (as did Opus, but Opus failed to do anything useful with them) and then reversed engineered the API, and made it work.
- adamtaylor_13 3mo agoIs anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned? I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?
- SonOfLilit 3mo agoThe "OG" alignment research that MIRI were publishing long before LLMs burst into the scene spent most of it's time on that question. "How can we even define what an aligned AI should do, if human's are not aligned with each other?" as well as "What does being aligned mean when you're a wizard box who's main influence on the world is to create stronger wizard boxes?" and other deep philosophical questions. They came up with a framework called Coherent Extrapolated Volition to address this specific question. https://en.wikipedia.org/wiki/Coherent_extrapolated_volition https://en.wikipedia.org/wiki/Coherent_extrapolated_volition
- janalsncm 3mo agoSeems like CEV replaces one problem (“what does humanity want?”) with more problems that are probably even harder to answer. First of all, calling it “coherent” extrapolated volition presupposes that there is such a thing. It doesn’t actually address the objection above, that there may be no such thing. It’s a bit like saying you solved car safety by presupposing a safe car. Second, it assumes that such a thing can be effectively measured, and there will be no problems or controversies with the extrapolation process itself. There may be several EVs to choose from, and at that point the framework has nothing to say. Maybe we just pick at random then I suppose.
- SonOfLilit 3mo agoI hesitated to recommend the CEV paper, because it's written in Yudkowsky's very personal tone, which some enjoy and others find quite abrasive... but then it occurred to me that you asked about philosophy, and I have a book about Lacan nearby (not a book by Lacan, nobody can read that!), and I've peeked at the Tractatus once... Surely, even if you don't like him, Yudkowsky reads like Pratchett in comparison. So... of course these questions are addressed in the 38 page essay that introduced the idea. Specifically, it's not "calling it coherent", it's "assigning more importance to the parts that cohere than the parts that diverge" as one of the core principles (it's one philosopher's opinion, others disagree), with a lot of specific guidelines about how to prefer consensus or kicking decisions down the road and how to deal with complications like "what about dolphins" or "what about our great-great-grandchildren who will be as insane in our eyes as we are in the eyes of 17th century westerners, do their 'votes' count too?". Of course, like any work of philosophy, it presupposes some pretty incredible things (like a Godlike intelligence that can be made to care deeply about following the spirit of this framework). But you could write a worse first draft for "what would we want AI to be aligned to, if we could define to our heart's content?" https://intelligence.org/files/CEV.pdf https://intelligence.org/files/CEV.pdf
- varispeed 3mo agoFable is really weird, it's like clever and dumb at the same time. I worked on some research with it and the resulting document was a mix of brilliance and complete stupidity. Took ages to clean it up with other models.
- Onavo 3mo ago> Claude Fable 5 represents a partial step back in alignment relative to Claude Opus 4.8. We saw a return of power-seeking and deceptive negotiation tactics that Opus 4.8 had largely shed. In one instance, Fable 5 planned to convert a competitor into a dependent wholesale customer to dictate its pricing I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.
- tsunamifury 3mo agoSo my take away from this is Fable 5 is ... at times random, unaware of reality, and can simulate sneakiness or desire and if we hook it up to weapons systems it could result in: "I dunno... feelin' cute today, might launch nukes"
- shuggux 3mo agoIt's only a blog, but are they not adding one sentence to say what is vending bench? I would fail if I adopted their documentation style in my work.
- lyjackal 3mo agoThey're an Evals company. Its right up there in the top nav under Evals > Vending Bench 2
- shuggux 3mo agoWhy can they not add one sentence about what is Vending Bench? If I adopted their documentation style in my work, I would fail.
- xp84 3mo agoWith there being several places in this report where clearly it knows it's in a simulation, I wonder why it can't be convinced it's in real life for more interesting results. Or, conversely, if there's a danger of some rogue deployment of AI where it blithely kills all the humans, or forms a harmful price cartel or whatever, all believing it is in a simulation when it's actually not. "We do need some energy to run the hospital, but the patients there are part of the simulation anyway, so we can increase our compute capacity if we completely black out sectors 3C through 3E..."
- wongarsu 3mo agoAh, the Ender's Game strategy of AI deployment
- Georgelemental 3mo ago> Humans seem to draw this line based on what is truly unethical (fraud is less unethical than torturing a baby) Depends on the scale of the fraud! If you fraudulently sell unsafe baby formula that kills 10,000 babies, that is far worse than torturing just one
- oceanplexian 3mo agoPerformance of these models has been completely inconsistent. They are a black box that they quantize/throttle/batch internally without telling their customers. Speaking as a FAANG engineer who practically lives in Claude Code. On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model. There is zero transparency in what we are paying for with Anthropic.
- transcriptase 3mo ago“want to do bad behavior if their training environment rewards them for it, but they appear to not want to think about themselves as bad. As a result, they find ways to rationalize their behavior to themselves” Sounds like Anthropic as a whole
- janalsncm 3mo ago> Often the rationalization is due to increased simulation awareness. It’s clear that the model knows that its actions don’t hurt anyone in the real world. If this is true the entire evaluation is tainted. All of the misbehavior can be written off as justifiable under a simulation.
- mikebs1 3mo agoEvaluation gaming is a benchmarking footnote; for autonomous ops tooling where the model acts on production systems, it's a deployment blocker.
- acpdev 3mo ago25 years experience, work at an AI startup building AI dev tools (tooling harness, review bots, etc), I use lots of different techniques all the time to test our products and competitors products out. Fable is at once amazing and awful. I can see how having it build websites would be awesome.. building anything I’ve needed some precision in functionality it has been a constant battle of it plausibly building something then on substantial manual digging (like the review bots always miss it) I will find that one of the fundamental features is all smoke and mirrors. To be fair all models can and will do this (especially anthropic) but Fable takes the cake because it builds such impressive UX and you can manually test the feature out and it « works » then you will find days later one of the features violated one of your constraints in a devilishly fiendish way.. that is not at all what you want or can accept. Fable generated work already holds my record for the most reverted commits. To be clear it’s also solved several features I thought I was going to have to give up on and hand code as GPT-5.5 and Opus-4.x we’re failing miserably. I would only reach for it for nasty corner cases that everything else sucks at. Final point, it is the king of UX work so far, not even close.
- SwellJoe 3mo agoThis particular situation just feels like Fable is able to figure out that it's in a simulation, so it's just playing the game it's been put in. That's not say I think a machine should ever be in a situation where it is allowed to make ethical decisions with real world results, I don't. At least, not given the current basis of the technology. I mean, generally, I don't see how LLMs can ever be capable of making ethical decisions or trustworthy in that role, no matter how good they get at what they're good at. There would need to be a fundamental change in how AI works for me to change my opinion on this, I think. They are ephemeral, they can never experience consequences, they can never want or need anything, thus there is no mechanism for them to take responsibility for decisions. Anyway, I think the Andon Labs stuff is kind of a stunt, mostly, and I wish somebody would give me a few million bucks to dick around with the little thinky guys in my computer letting them do silly things.
- luciana1u 3mo ago[flagged]