4 ms·
This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. Fr
by redfloatplane 4mo ago
This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. From the system card:
> 6.2.5 External testing from Andon Labs
Andon Labs reviewed the behavior of Claude Opus 4.8 in their simulated Vending-Bench 2
retail-management evaluation, as reported in the Capabilities section of this system card
(see Section 8.13.5). Although they did observe some unexpected capability failures, they
did not find clear instances of the kind of concerning in-game behaviors that were
discussed in other recent system cards.
> What might have led to these differences? We monitor and investigate the effects of
different training environments on alignment; Claude Opus 4.7, for example, had training
that focused on business skills and robustness against adversarial agents, but we
discovered that this training inadvertently contributed to misaligned behavior including
dishonesty. We therefore removed it for Opus 4.8.
> Thus, Opus 4.8 did not show the same misaligned behaviors as Opus 4.7 in Vending-Bench,
but also had reduced business success due to being more susceptible to scammers and
being less able to negotiate good deals with other agents. We are currently working on
training to improve business capabilities while maintaining aligned and ethical behavior.
- mrdependable 4mo agoI don't know how people can read stuff like this and think LLMs are intelligent or conscious.
- stratos123 4mo agoConsciousness aside, why does reading about an LLM generalizing from specific to general dishonesty make you think it's not intelligent?
- redfloatplane 4mo agoI don't really see how you got to your comment from what I quoted. However, somewhat relatedly, I proposed a thought experiment about this in the comments for Opus 4.7[0]: > It's April, 1991. Magically, some interface to Claude materialises in London. Do you think most people would think it was a sentient life form? How much do you think the interface matters - what if it looks like an android, or like a horse, or like a large bug, or a keyboard on wheels? > I don't come down particularly hard on either side of the model sapience discussion, but I don't think dismissing either direction out of hand is the right call. [0]: https://news.ycombinator.com/item?id=47680059 https://news.ycombinator.com/item?id=47680059
- mrdependable 4mo agoWith the amount of data these models have, they should be much more capable if there was an actual intelligence behind it. If you saw someone running into a wall continuously until you showed them how to use a door, even though they have seen people use doors a million times, what would you call that? The fact that Anthropic needs to poke, prod, and guide these models to behave in the desired way does not give the impression of intelligence. It gives the impression of a complicated automaton.
- solenoid0937 4mo agoThis is such a bizarre statement, you speak as if you have any understanding of how much data "should be" required to make an intelligence but frankly you don't know. None of us know.
- mrdependable 4mo agoI am not talking about how much data is required to make intelligence. I am talking about how it uses the data it already has. It can tell you about every scam in the book, research about the scams, how to spot scams, who does the scamming, etc. Everything under the sun about scams. However, without the “skill” included in a prompt it will fall for scams.
- Topfi 4mo agoIf I may slightly tweak your example to highlight why I find it very flawed: > It's April, 1891. Magically, a drone swarm with lights piloted to show a face [0] materialises above London. Hidden speakers command the public to listen, for this is their Gods arrival. Do you think most people would think this was a religious entity? What if the drone pilots decided to adjust to something the local populous would expect to see during the second coming, does that matter? We cannot, nor should we discard what we know about LLMs and their limitations. Such examples are not really helpful and it is very reductive to take the "walks like a duck" approach to autoregressive models in 2026, when we have ample evidence that these, while powerful and capable in a lot of use cases, are not in any way comparable to actual reasoning. With EBM [1] we already have empirical evidence that other solutions can get us closer to actual artificial reasoning (though whether these get us fully there remains to be seen, I tend to lean on "extraordinary evidence" for any such statement at this stage). [0] https://www.youtube.com/watch?v=YH1BD7kKqKw https://www.youtube.com/watch?v=YH1BD7kKqKw and of course https://www.youtube.com/watch?v=dy2zB8bLSpk https://www.youtube.com/watch?v=dy2zB8bLSpk [1] https://logicalintelligence.com/blog/energy-based-model-sudoku-demo https://logicalintelligence.com/blog/energy-based-model-sudo...
- asdewqqwer 4mo agoAs if the dishonesty of human who are good at business has not been criticized since business ever exists
- kurtoid 4mo agoThe H in business stands for honesty
- bbbbread 4mo ago[dead]
- bbbbread 4mo ago[dead]