14 ms·
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
by afavour 2mo ago
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
- fl4regun 2mo agoi don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
- afavour 2mo agoDestined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
- Matl 2mo agoI agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating. Granted, this can probably be tuned for.
- afavour 2mo agoAnd really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.
- horsawlarway 2mo agoIf you read the full post, I'm not actually sure I agree with the title. Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior. To recap: 1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself. 2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers. 3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours". Frankly... I'm more annoyed at the posters than the bot.
- rcxdude 2mo ago> “Of course the AI lied and cheated, the task it was given was really difficult!” It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
- blargey 2mo agoFail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well. “Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
- fl4regun 2mo agomaybe it is because I am biased but I have almost no expectation for AI to act "ethically"
- jerf 2mo agoDo you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
- butlike 2mo agoNo matter the urgency, you shouldn't sacrifice your ideals. That's why they pay you; to fall on the knife
- JohnMakin 2mo agoThey aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test. Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all. In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery. And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
- 2mo ago
- RHSeeger 2mo ago> Results that arrive after the deadline do not exist Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)
- throwatdem12311 2mo agoSounds like every startup I ever worked for. What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever
- infinite_spin 2mo ago> What’s the line? Evidence, even when downplayed or ignored, is still evidence.
- bpodgursky 2mo agoHumans care about reputation and legal repercussions from fraud, that persist after business failure. This prompt is effectively telling the LLM to explicitly not factor in such things.
- queenkjuul 2mo agoTraining data imparts that desperate people lie, but not that lying has consequences?
- Maxatar 2mo agoYou might be reading the prompt far too literally then. LLMs interpret words not based on literal and rigorous definitions but based on how those words are actually used in reality based on a large corpus of text. In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have. The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.