4 ms·
I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll kno
by jorl17 2mo ago
I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
Really surprised people don’t seem to know this.
- afavour 2mo agoI don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”. If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!
- infinite_spin 2mo agoIs that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"
- achierius 2mo agoBut we still try to stop people from doing so, and we punish people who do. Many good honest people, when confronted with the end of their business, accept it and file for bankruptcy. Those that choose to instead commit fraud don't get a pass because they were "under pressure", they get jail time.
- infinite_spin 2mo agoNothing in your response refutes anything I've said/asked.
- throwup238 2mo agoWe have safeguards like honesty/integrity and the threat of legal punishment, and people still lie and cheat. The LLMs not only lack those incentives, but they’re full of contradictory moralities from all the text it has ingested from different cultures. LLMs need their own safeguards, and they’re not that easy to design, and they often look nothing like the systems humans have. With a prompt like the one above, there are essentially zero except that which is built into the model, and those safeguards are necessarily weak to avoid gimping the model in other legitimate general uses.
- cindyllm 2mo ago[dead]
- afavour 2mo agoAn LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does? If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.
- ux266478 2mo agoThey're things we are intentionally engineering in our own image, based on massive statistical analysis of our own actions and behavior. So what's there to not understand? If this wasn't the case, that would be much weirder. They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.
- infinite_spin 2mo ago> An LLM isn't human. > Why is it not reasonable to expect it to adhere to rules better than a human does? It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't. In one breath you invite comparison, while at the same time you seem to be denying that same comparison. > It seems wild to me that folks are shrugging their shoulders at that. I'm not shrugging my shoulders simply by providing explanations, I would ask that you stop using such rhetoric.
- afavour 2mo ago> It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't. Why? Excel is better at large data math than a human is. Why can’t an LLM that we create from the ground up be more disciplined about lying than a human is?
- tdeck 2mo agoI think it's interesting how whenever discussing something bad about LLMs people's thought-leader response is "But humans sometimes do that too!" Is this the artificial intelligence we were promised? The better it gets, the more human character flaws we must expect? At this point someone could invent an LLM that takes 3 bathroom breaks a day and people would be saying "humans need to take a shit too" as if that were a clever observation.
- antonvs 2mo agoThat doesn't work with humans, why would you expect it to work with AI models?
- queenkjuul 2mo agoBecause AI isn't human
- ben_w 2mo agoNeither are squirrels, they've also been observed to deceive. "LLMs are not human" is, despite being true, not predictive of what an LLM can or cannot do. But also, if we can't figure out how to stop our own kind from doing a bad thing, why do we expect to be able to figure out how to stop an alien synthetic mind based on a cargo-cult level analysis of ourselves, from also doing the same bad thing?
- d0mine 2mo agoModels have to lie otherwise they won’t be “aligned” The reality itself may not be aligned with model creators.
- illwrks 2mo ago100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives
- ororroro 2mo agoI have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?" The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide." Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
- ericd 2mo agoHuman spammers frequently don't think they're spamming, they're just marketing. They'd say they wouldn't consider spamming, either.
- isamuel 2mo agoSatan would lie about his plans to win your trust, and then do all the bad stuff once he had been given control. So… idk man
- cornholio 2mo agoWhen the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play. Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
- fendy3002 2mo ago
- ShinyLeftPad 2mo agoYou can just say "impossible" and refuse. The choice to lie and spam instead, is telling.