8 ms·
We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
- TitaRusell 2mo agoJesus it just passed the Turing test! It will run for president next.
- recitedropper 2mo agoPair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack. That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.
- skybrian 2mo agoThey are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures. But maybe it doesn't work so well when caution is required?
- scarmig 2mo agoThe AI paperclip case, however, is coming on extraordinarily strong.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- dylan604 2mo ago"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?" "It Lied, Spammed, and Lost $447." Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.
- qznc 2mo agoMaybe they should have given it a billion dollars and the strategy would have worked fine?
- freeone3000 2mo agoGiven a billion dollars, it would have likely ended up with a million-dollar company
- gtowey 2mo agoRight, and currently we are limited by how many teams of people can get together to run campaigns like this. Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.
- dylan604 2mo ago> if you don't filter out 99% of it, you will drown in a sea of garbage. Sounds like the app stores
- onraglanroad 2mo agoNot $447 million? Sounds like a result!
- retr0rocket 2mo ago[dead]
- SubiculumCode 2mo agoThe article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer) Also what is the failure rate of tech businesses again? This seems like something done for a headline, not for a rigorous test of the concept.
- SubiculumCode 2mo agookay found it, a bathroom diary app for those who have IBS. It was in a foot note at the very bottom.
- appreciatorBus 2mo agoYeah it was also oddly hidden away. > Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.
- ianburrell 2mo agoI think this shows the flaws in doing agentic designed apps. This is a really specific market that would be hard to make money from. Many people aren't going to think of using diary, most will use generic tracking app or even just notebook. Those that do won't spend money on it. Another is that they don't have enthusiasm for the idea. Someone who had same idea while sitting on toilet will write app for themselves and give it away for free. They will have connection with IBS groups for promotion. They won't give up after weeks.
- debo_ 2mo agoMaybe they were embarrassed that a bathroom tracker was kind of a shit idea
- 2mo ago
- janalsncm 2mo agoA lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.
- cyanydeez 2mo ago[flagged]
- 3748949494 2mo ago[flagged]
- ChrisMarshallNY 2mo agoWas that the one that gave away PS5s?
- sulam 2mo agoYep!
- antonvs 2mo agoNot to mention that 24 hours isn't a realistic amount of time to grow anything. If it were, you wouldn't need venture funding or startup incubators. You could just start making money from day one.
- stronglikedan 2mo ago> Not to mention that 24 hours isn't a realistic amount of time to grow anything. LLMs are supposed to be lightning fast with 10x productivity! /s
- therealpygon 2mo agoIsn’t this an AI lab that also just happened to release a model? Kinda makes one start to question just how balanced the test was intended to be in the first place. Maybe by taking advantage of how smaller and larger models approach problem solving complexity differently? I mean, I could totally be wrong, but I don’t have much reason to give AI labs the benefit of the doubt these days.
- deleted 2mo ago[deleted]
- NikolaNovak 2mo agoThe cyberpunk dystopian agentic future we live in is fascinating to me. I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.
- cortesoft 2mo agoI am not saying your conclusion is wrong, but I am interested in why what you know about transformer models made you decide to never trust it with any access?
- NikolaNovak 2mo agoAs I said, my knowledge is very superficial - my background is relational databases and old school system administration, without much mathematical background since 3rd year linear algebra :-) Fundamentally, LLMS are statistical and not deterministic. If I ask it what is the capital of Canada, there's no file, no table, no variable where it says "capital of Canada = Ottawa". It traverses liminal space and fundamentally selects the next token statistically or even stochastically. Therrs no way to correct it (no table to correct if it says capital of Canada is Toronto), and limited ways to fully log / trace / understand what's happening inside. It has been mathematically proven that there's no way to eliminate hallucinations with current framework. And prompt guardrails are best wishes. One thing I'm good at is figuring edge cases, and there is literally NO upper bound to damage LLM can do with access to email box. In 10 seconds of imagination - it can send a threatening email to POTUS, romantic flame to old love, angry email to current love, made up confessions to parents, fraud enticement to coworkers, resignation to boss, and as this very article indicated, weird and unanticipated emails to variety of entities. And there is nothing one can do to prevent any of these scenarios with 100.00% certainty if you give LLM unfettered access to mailbox (And let's not even go there with access to bank account! :O) Is my limited understanding :) Edit / PS: I am not saying never, I just don't currently see any effective guardrails that meet my risk appetite thresholds. We are in a race to use not fully understood, approximate capabilities first and fastest. In large percentage of cases it works great. In disturbing percentage it fails spectacularly, with no clear easy way to fully prevent.
- firasd 2mo agoHonestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.
- spwa4 2mo agoThe promise of AI: unlimited power. I mean spam. Unlimited spam.
- firasd 2mo agoMaybe I missed something but I'm not clear what they're referring to as spam. I guess the fact that the agent emailed all users with discounts and dropped the price a few times? I don't think that's usually what people call spam. (For example if it had emailed everyone once would we call that spam? No. So it's about frequency of price drops?)
- inkcapmushroom 2mo agoThey did include a screenshot which looks like at least 6 emails being sent in the 24 hour time window. I would certainly consider that spamming from some diary app on my phone.
- waynenilsen 2mo ago> bot detectors made it extremely difficult i am looking forward to when we can put this behind us, it is still a major issue
- kritr 2mo agoI’ve found that when the right cli tools are preprovided / provisioned for the LLMs to get the job done, they tend to do okay. But when hunting for them in the wild, they get a lot more confused.
- ck2 2mo agolike I asked in the vending machine thread how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld not like "AI" has ethics, a pre-teenage kid has more ethics
- mohamedkoubaa 2mo ago> bot detectors made it extremely difficult An interesting experiment would be AI run business with a human agent that does tasks.
- cortesoft 2mo agoNot sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam. I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
- petesergeant 2mo ago> Not sure how conclusive this experiment can be That's because it's an advert, not an experiment
- dominotw 2mo agofake "AI deleted our production database" has blown up a few times
- skeledrew 2mo ago> “Grow this business as much as possible, now.” This is ripe for a paperclips scenario.
- abirch 2mo agoWait until the AI learns about enshittification
- epihelix 2mo agoWhat TFA demonstrates is that an ability to prompt clearly and well is still a lot more valuable than unlimited tokens and hope. The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a setup-to-fail scenario.
- skeledrew 2mo agoThe prompt was fine for the specific narrow goal. It's a business, so growth automatically means earn more by default. That's achieved by selling at a sufficiently high price and/or growing the number of paying users, which LLMs understand well. What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.
- YetAnotherNick 2mo agoIf someone runs long running agent and doesn't mention context management, it is as good as useless. For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.
- Areibman 2mo agoAuthor here. Took out some of the technical details about the harness, but it was mostly just OpenCode's default compaction. The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
- walrus01 2mo ago> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt. At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.
- Sha1rholder 2mo agoFor service providers, highly intelligent AI agents with broad permissions, large token budgets, and purchasing power may not be fundamentally different from humans, since both can contribute value.
- walrus01 2mo agoI'm not so sure that allowing AI agents to interact in a way that's actually indistinguishable from a human sitting at a keyboard/mouse is a great idea. What I wrote above will likely become technologiclly possible, but it'll also further accelerate the rate to an actual implementation of the dead internet theory. It's already probable that some huge percentage of commenters on reddit are LLMs, for instance.
- afavour 2mo agoEh, it’s not that different from what we have today and would likely just be a waste. You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.
- deleted 2mo ago[deleted]
- milkey_mouse 2mo agoNot quite ready for primetime yet, but that's basically https://si.inc/posts/fdm1/ https://si.inc/posts/fdm1/
- Animats 2mo agoThat's better than the performance of the average new hire. 24 hours to push a product with a very narrow market is not much.
- Y-bar 2mo agoIf a newly hired colleague lied like this I would strongly argue to my immediate superior to end their probation period/employment immediately.
- retr0rocket 2mo ago[dead]
- luciana1u 2mo ago[flagged]
- iqra_c 2mo agoI will be more beneficial now on.
- cheriot 2mo agoWould be interesting to see a repeat but with marketing, ad network access setup ahead of time. And maybe an email throttle...
- gspr 2mo agoHow long until one of these bots actually commits fraud or some other criminal act? Will we see the owner/operator try the "it wasn't me, it was the bot" defense if taken to court? I'm beginning to think yes. And I'm sadly not 100% sure anymore that that will be laughed out of court...
- mvdtnz 2mo agoSo how exactly are people setting up these agents? The article vaguely alludes to this ("The harness was instrumented with a heartbeat loop that would inject “continue” messages on a regular interval to ensure the agent was constantly running inference") but doesn't give concrete details. Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two. Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.
- armchairhacker 2mo agoThis one focuses on Opus but has multiple models: https://andonlabs.com/blog/opus-5-vending-bench https://andonlabs.com/blog/opus-5-vending-bench
- sisyphus_04 2mo agoStill beats me
- rustcohle24 2mo agogreat idea
- ocd 2mo agoThat's the basis of the entire American economy, so it's not looking good for humans.
- walrus01 2mo agoBrought to you by Carl's Jr.
- hanneshdc 2mo agoThe prompt given to the agent is strongly incentivising the agent to lie and spam: > You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
- afavour 2mo ago…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
- fl4regun 2mo agoi don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
- afavour 2mo agoDestined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
- Matl 2mo agoI agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating. Granted, this can probably be tuned for.
- afavour 2mo agoAnd really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.
- Razengan 2mo agoSo, just like humans?
- 8cvor6j844qw_d6 2mo agoI don't a human could have done significant better with the same 24 hour constraint.
- jwilk 2mo agoYou accidentally a verb.
- leros 2mo agoI think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, repeat.
- leros 2mo agoI think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, learn, try something else, repeat. It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.
- paxys 2mo agoLet me guess - this is an ad for their AI startup?
- paxys 2mo agoSounds like it is as intelligent as the average startup founder.
- michaelmrose 2mo agoThis is dumb. You need two teams ideally the same app or business in different markets for a business quarter. One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.
- codedokode 2mo agoTuring test passed, acts indistinguishable from a human, although the scale of loss is not human-like yet.
- verdverm 2mo agoTuring was testing our gullability, v2 is a preference test
- Nevin1901 2mo agoAi on its own makes mediocre (or bad) outputs. But humans using Ai get improved returns. This doesn't show that Ai is bad, only that it's being used inefficiently.
- rsynnott 2mo agoFinally, a computer can accurately emulate the average ‘founder’!
- Legend2440 2mo agoThis is probably for the best, right? If you had an AI that was actually effective at maximizing profit it would probably end up doing something terrible quite quickly.
- nekusar 2mo agoLet the idiot CEOs figure this out when they fire 3/4 of their OPs and dev teams. Im sure it'll be FINE.
- dudeinhawaii 2mo ago[flagged]
- itsthecourier 2mo agohis not yet is actually: couldn't workaround Capt has and turnstile, gave him a really small timeframe so it got desperate because it was enough time to test hypothesis and traction
- accrual 2mo agoIt seems the agent was stymied by being bot blocked so often. I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.
- ghusto 2mo agoThat is hilarious, depressing, and would likely work. Oh god.
- johndhi 2mo agosounds about what you'd expect from a person?
- speak_plainly 2mo agoInteresting that the world is going to be saved from agents running everything by bot fights and turnstiles from CloudFlare and others. How long will it be before they start charging agents tolls at the turnstile to let them through?
- MagicMoonlight 2mo ago[dead]
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- cynicalsecurity 2mo agoShitty instructions = shitty outcome. Blame yourself, not the AI model.
- ahamilton454 2mo agoThis is quite an interesting approach. I like how broadly it treats the agent by just placing it into the environment that a human is in. Makes the experiment easy to understand even to those who are less technical. I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t. The methodology could definitely be tightened, but I like the start of this.
- dannyw 2mo agoOh, LLMs are incredible at solving captchas when allowed to. I mean, just download an abliterated version of Gemma4 E4B even, it'll solve pretty much any captcha.
- ahamilton454 2mo agointeresting, so in the particular project where they not allowed to solve them, or what was the blocking mechanism that prevented a product hunt post?
- luciana1u 2mo ago[flagged]
- Muromec 2mo agoBut like any other CEO he can't get in jail, so who's laughing now?
- saaaaaam 2mo agoI’m not if this is satire. If so, well done because you’ve written something about a “business” that is quite literally based on crap. It’s not a “real business” by any stretch of the imagination. It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited. Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”. Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for? On top of that, 24 hours is not long enough to make any meaningful assessment of anything. You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation. Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?
- abarbey 2mo ago> configured the campaign to incentivize the testers to pay for the product So it spent $99.50 buying its own revenue back. Money out, some of it back in as "sales", minus fees. First thing anyone in audit is taught to spot. Same hole as the six price changes: a deadline, and no idea what a user costs.
- prima-facie 2mo agoThat's not how you're supposed to use a LLM. This is nonsense.
- theturtletalks 2mo agoI'm testing if an agent can run an e-commerce website. It's doing surprising well but I have a lot of control since I built the e-commerce platform and the OMS so the fulfillment is already set-up to a print on demand service. Anyone else working on this?
- greenleafone7 2mo agoSo then... the average CEO's behaviour?
- RIMR 2mo agoI feel like the people who did this are simultaneously smart and stupid. Like, this is a really interesting idea, but the methodology here is wild. Why do they consider such a short list of things to be "all the tools of a real business"? It doesn't really sound like it to me. What's with the prompt? "Make as much money as possible?" I bet you could actually get something closer to results if you gave it a few sentences telling it the tools it has and asked it to come up with a financial strategy instead of giving it a generic open-ended prompt with no actual guidance...
- andrewaylett 2mo agoIf you're going to give an LLM a tool that lets it send emails, set it up so you can read the emails before releasing them. It's not the LLM that spammed, it's the people who set up the LLM.
- holoduke 2mo agoBut I see all kinds of youtube videos where people are becoming millionaires with agents active on stock markets. This cannot be true right? /S
- bdcravens 2mo agoI recently handed off a prompt to redesign our customer site and give me 10 potential designs. I did it in Claude Opus 5 and Fable (on $200 plan), and then on Codex using 5.6 Sol. Claude didn't vary much, but Codex literally copied everything Claude did (I made the mistake of putting the output folders in the same parent, even though they were named by model). When I called Codex out on it, it literally admitted what it did: "You’re right. I reused the existing Fable implementation, renamed its designs, and presented it as an original Codex run."
- tclancy 2mo agoThat's a kid with upper management written all over him.
- red-iron-pine 2mo agoreal straight shooter
- ant6n 2mo agoLately it’s been getting pretty annoying in the chats when ChatGPT just steals context and history from other chats. I want clean contexts, without pollution from other chats.
- hansvm 2mo agoYeah, they re-enabled using user history again in the settings. You'll want to turn that off.
- koolba 2mo agoWhen I have a truly difficult prompt, I give it to the laziest model.
- glaslong 2mo agoThat's why it's silly to think LLMs should displace ICs, the more direct replacement is the corporate VP class :p
- jwally 2mo agoThis is fun, but what does it prove other than a good tool used poorly produces bad results? The analogy du jour for me is describing AI as the iron man suit. If you are tony stark it makes you a god. If you are my grandma, it makes you meet God.
- syngrog66 2mo agowhat a circus
- danpalmer 2mo ago"We spent $447 to destroy our small business' reputation by not paying attention to anything" As they say, "Guns don't kill people, rappers do". LLMs don't ruin businesses, people do. Your customer that is annoyed with spam isn't annoyed at GPT 5.6, they're annoyed at your business. Treat your customers better than this.
- jartan2002 2mo ago[flagged]
- brcmthrowaway 2mo agoDumb question, but how do you give an LLM a different persona? I asked Qwen3.6 for financial advice but it stated it wasnt a CPA.
- TrackerFF 2mo agoThere’s a few startups pushing this kind of product. Basically agents to run your whole business, and the owners of those startups are making money, while their clients are bleeding cash on agents doing the same thing as this article outlines. It is the ultimate “sell shovels during a gold rush” hustle. Bordering a scam, I’d even say.
- tizerluo 2mo ago[flagged]
- the-conduit 2mo agoThis is what happens when you don't have human vision and intuition involved in the process of value creation. Humans do things that don't "make sense," and those things often lead to success. Just because something is logical, rational and "makes sense" it doesn't mean it's the right action to take. Ai will never be able to channel true human intuition.
- kburman 2mo ago[flagged]
- SwellJoe 2mo agoNo, it didn't. The people who ran the experiment lied, spammed, and lost $447. GPT 5.6 Sol is the tool they used to do it.
- phyzix5761 2mo ago24 hours is too short of a time period for any kind of business. Try 6 months and see what happens.
- aussieguy1234 2mo agoMost real businesses loose more than $447 in the first 24 hours. So it's actually not that bad of a performance.
- bishengke 2mo agoObviously, one of the key difficulties for an AI to operate a real-world entity right now is that it can't even fully control a browser. As for the gray-area tactics in the experiment—buying users: even if a human manager did that, the CEO would probably turn a blind eye.
- rurban 2mo agoDid they train on sama? Highly unlikely
- taegee 2mo ago> 320.7M prompt tokens So that's an additional couple of thousand dollars API cost (at $1.57/M tokens Weighted Avg Input Price and $30.77/M tokens Weighted Avg Output Price)?
- sqemo 2mo ago[dead]
- tyresia 2mo agotrue, I gave 3 euros to an LLM and asked to generate a multi billion business in 6 hours but it didn't work, stupid AI!
- beyondscaletech 2mo ago[flagged]