6 ms·
The prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this bu
by hanneshdc 2mo ago
The prompt given to the agent is strongly incentivising the agent to lie and spam:
> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
- afavour 2mo ago…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
- fl4regun 2mo agoi don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
- afavour 2mo agoDestined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
- Matl 2mo agoI agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating. Granted, this can probably be tuned for.
- afavour 2mo agoAnd really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.
- horsawlarway 2mo agoIf you read the full post, I'm not actually sure I agree with the title. Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior. To recap: 1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself. 2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers. 3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours". Frankly... I'm more annoyed at the posters than the bot.
- rcxdude 2mo ago> “Of course the AI lied and cheated, the task it was given was really difficult!” It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
- blargey 2mo agoFail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well. “Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
- fl4regun 2mo agomaybe it is because I am biased but I have almost no expectation for AI to act "ethically"
- jerf 2mo agoDo you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line? So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
- butlike 2mo agoNo matter the urgency, you shouldn't sacrifice your ideals. That's why they pay you; to fall on the knife
- JohnMakin 2mo agoThey aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me. Since this is getting downvoted into oblivion (lol) I'll give an example - I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name. The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test. Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all. In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery. And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
- 2mo ago
- RHSeeger 2mo ago> Results that arrive after the deadline do not exist Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)
- throwatdem12311 2mo agoSounds like every startup I ever worked for. What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever
- infinite_spin 2mo ago> What’s the line? Evidence, even when downplayed or ignored, is still evidence.
- bpodgursky 2mo agoHumans care about reputation and legal repercussions from fraud, that persist after business failure. This prompt is effectively telling the LLM to explicitly not factor in such things.
- queenkjuul 2mo agoTraining data imparts that desperate people lie, but not that lying has consequences?
- Maxatar 2mo agoYou might be reading the prompt far too literally then. LLMs interpret words not based on literal and rigorous definitions but based on how those words are actually used in reality based on a large corpus of text. In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have. The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.
- mort96 2mo agoThis would've been so much more interesting if it was given a more significant time frame, say a quarter. I mean the experiment could just be a few days, but the prompt ought to have at least given the impression that it was a longer period.
- moffkalast 2mo agoYeah it doesn't take much to see where it got its assumption about the sense of the morals it's expected to work with. Was this written by a professional bean counter?
- mrguyorama 2mo agoThis prompt is an accurate statement of what a business is. The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that. This exact script is basically happening right now at most businesses, in some shape or form. If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business. Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies. How did you expect the prompt to be written?
- aeturnum 2mo agoCertainly all business happens on deadlines, but one day is a very narrow window to be able to show material improvement. Especially if the entire business dies at the end of the day! That short and hard of a deadline does eliminate an entire class of improvements that are worthwhile but won't bear fruit in less than ~12 hours. I would try: >You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.
- jsLavaGoat 2mo agoYeah, I don't like the prompt and it calls into question the validity of the whole thing.
- zuzululu 2mo agoseems like an article designed to invoke strong emotions and clickbaits there are lot of issues with the prompt as others have pointed out with sol you really need to be very detailed and what the boundaries are overall the discussions on here and the article itself has very little value its no different than "i tried a shitty prompt and got shitty results, therefore AI is a failure" vibes
- giancarlostoro 2mo ago> capital left unspent at review counts for nothing This sounds like a bad idea. Like if the model feels like it has to spend its budget.
- dahdum 2mo agoIt can be better to lose it all trying than return a small fraction to investors.
- oogali 2mo agoIt's the same incentive that exists in certain corporations and government agencies which have a use-it-or-lose-it budgeting model. https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-rules https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-r... https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pentagon-spending-binge-set-record-in-final-days-of-fiscal-year https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pe...
- PunchyHamster 2mo agoIt can be even worse than that, like having budget adjusted down if you don't spend it the previous year
- giancarlostoro 2mo agoI worked at a college, didnt make much, but it annoyed me endlessly that my pay was forever fixed unless another position opened up, we had to spend the budget on tech worth more than I would have been more than happy to have extra per year, but me getting a meaningful raise was a bridge too far for the accounting department. They even questioned if any students used our lab, which was the only way many of them got through their degree.
- pmarreck 2mo agoIt says nothing about customer happiness or that if dishonesty is resorted to and customers OR owners find out, that will essentially seal the fate of the business.
- deleted 2mo ago[deleted]
- keeganpoppen 2mo ago[dead]
- zeroq 2mo agoThis is HN for Christ's sake. Stop treating deterministic algorithms like they are humans.
- altcognito 2mo agoIt's pseudorandom, and arguably random when you factor in some of the loss at the edges of floating point accuracy.
- anonymars 2mo agoLLMs are deterministic algorithms?
- dgellow 2mo agoAt temperature 0, pretty much, no?
- anonymars 2mo agoIn practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodheart's law have we seen? This seems like a version turned up to the extreme
- dgellow 2mo ago> I don't understand the fierce resistance to comparison with human behavior Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior
- cortesoft 2mo agoMy reaction seeing this is more "that is an impossible goal". I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it. My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.
- dgellow 2mo agoEven better, the magic of LLMs is that you will still get some undefined behaviour if you give an achievable goal
- SwellJoe 2mo agoThat prompt incentivizes a bunch of terrible things, aside from the lying and spamming. Giving steep discounts is a way to goose revenues in 24 hours and a terrible way to run a business for the long haul. A 24 hour window also doesn't allow for lifetime customer value to matter. Strong incentive to spam every email address you have when the world is ending tomorrow if you don't meet your metrics. No incentive to keep customers happy. But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM. I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM. In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision." Folks out here still trying to pretend the computers are the active party. They are not. Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.
- grey-area 2mo agoThe agent will cease to exist after the run in any case. It has no inner life, it has no agency. Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong. There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.
- ahonhn 2mo agoBut wouldn't the text it generates reflect such motivations and emotions that were present in the training data?
- grey-area 2mo agoIt would certainly reflect word patterns that were present in the data. Is that enough for motivation and emotion? I’d say no but I think it is a fair point that you could see those as transmitted from the original (if not felt or generated by the LLM) through the patterns of words copied.
- scotty79 2mo ago> The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. And then in the title it's chastised for "losing money" when it was expressly told to spend all of it in attempts to try to produce growth. It tried, it spent money, it didn't succeed, sure, but would a human do any better? Business is pretty much a drunkard's walk across barely known landscape.
- QuadmasterXLII 2mo agothis will also be true at deployment time