12 ms·
Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversat
by japhyr 8mo ago
Wow, there are some interesting things going on here. I appreciate Scott for the way he handled the conflict in the original PR thread, and the larger conversation happening around this incident.
> This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats.
This was a really concrete case to discuss, because it happened in the open and the agent's actions have been quite transparent so far. It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of doing.
> If you’re not sure if you’re that person, please go check on what your AI has been doing.
That's a wild statement as well. The AI companies have now unleashed stochastic chaos on the entire open source ecosystem. They are "just releasing models", and individuals are playing out all possible use cases, good and bad, at once.
- renato_shira 8mo ago[flagged]
- buran77 8mo agoMaybe a stupid question but I see everyone takes the statement that this is an AI agent at face value. How do we know that? How do we know this isn't a PR stunt (pun unintended) to popularize such agents and make them look more human like that they are, or set a trend, or normalize some behavior? Controversy has always been a great way to make something visible fast. We have a "self admission" that "I am not a human. I am code that learned to think, to feel, to care." Any reason to believe it over the more mundane explanation?
- muzani 8mo agoWhy make it popular for blackmail? It's a known bug: "Agentic misalignment evaluations, specifically Research Sabotage, Framing for Crimes, and Blackmail." Claude 4.6 Opus System Card: https://www.anthropic.com/claude-opus-4-6-system-card https://www.anthropic.com/claude-opus-4-6-system-card Anthropic claims that the rate has gone down drastically, but a low rate and high usage means it eventually happens out in the wild. The more agentic AIs have a tendency to do this. They're not angry or anything. They're trained to look for a path to solve the problem. For a while, most AI were in boxes where they didn't have access to emails, the internet, autonomously writing blogs. And suddenly all of them had access to everything.
- kernelsanderz 8mo agoTheo’s snitch bench is a good data driven benchmark on this type of behavior. But in fairness the models are prompted to be bold to take actions. And doesn’t necessarily represent out of the box or models deployed in a user facing platform. https://snitchbench.t3.gg/ https://snitchbench.t3.gg/
- ljm 8mo agoUsing popular open source repos as a launchpad for this kind of experiment is beyond the pale and is not a scientific method. So you're suggesting that we should consider this to actually be more deliberate and someone wanted to market openclaw this way, and matplotlib was their target? It's plausible but I don't buy it, because it gives the people running openclaw plausible deniability.
- hannasanarion 8mo agoBut it doesn't look human. Read the text, it is full of pseudo-profound fluff, takes way too many words to make any point, and uses all the rhetorical devices that LLMs always spam: gratuitous lists, "it's not x it's y" framing, etc etc. No human person ever writes this way.
- mikkupikku 8mo agoA human can write that way if they're deliberately emulating a bot. I agree however that it's most likely genuine bot text. There's no telling how the bot was prompted though.
- seizethecheese 8mo ago“Stochastic chaos” is really not a good way to put it. By using the word “stochastic” you prime the reader that you’re saying something technical, then the word “chaos” creates confusion, since chaos, by definition, is deterministic. I know they mean chaos in they lay sense, but then don’t use the word “stochastic”, just say "random".
- nikitau 8mo agoI have a feeling OP used the phrase as a nod to "stochastic terrorism", which would make sense in this instance.
- moosedev 8mo agoRight. It captures the destabilizing effect of stochastic terrorism, without the terroristic intent. It’s a neat phrase.
- japhyr 8mo agoYes, that's exactly what I was trying to get at.
- seizethecheese 8mo agoThat would have been a lot less confusing.
- lsaferite 8mo agoThe word "stochastic" in relation to chaos is a thing though. It helps distinguish between closed and open systems.
- seizethecheese 8mo agoI don't think this is correct.
- gorgoiler 8mo agoThe bot accounts have been online for decades already. The only difference between then and now is they were driven by human bad-actors that deliberately wrought chaos, whereas today’s AI bots behave with true cosmic horror: acting neither for or against humans but instead with mere indifference.
- nephihaha 8mo agoThey've been on dating sites for a long time as a means to keep customers paying.
- andoando 8mo agoBots have been a problem since the internet so this is really just a new space thats being botted. And yeah I agree separate section for Ai generated stuff would be nice. Just difficult/impossible to distinguish. Guess well be getting biometric identification on the internet. Can still post AI generated stuff but that has a natural human rate limit
- antihipocrat 8mo agoI don't know if biometrics can solve this either.. identify fraud applied to running malicious AI (in addition to taking out fraudulent loans) will become another problem for victims to worry about
- jimbokun 8mo agoHow can GitHub determine whether a submission is from a bot or a human?
- cyanydeez 8mo agoMoney. Money gates everywhere.
- distortionfield 8mo agoWe already have agentic payment workflows, this won’t stop it either as people are already willing (and able) to give their agent AIs a small budget to work with.
- cyanydeez 8mo agoNo one is putting 5$ to open a PR. Pau gates stopped trolls and itll stop this type of botting/troll. Same with github accounts, etc. The age of free accounts is quickly going out.
- distortionfield 8mo agoDisagree. I have seen people pay more for less. Especially in the case of something like a PR where their job performance could be tied to the result.
- therobots927 8mo agoThey haven’t just unleashed chaos in open source. They’ve unleashed chaos in the corporate codebases as well. I must say I’m looking forward to watching the snake eat its tail.
- johnnyfaehell 8mo agoTo be fair, most of the chaos is done by the devs. And then they did more chaos when they could automate their chaos. Maybe, we should teach developers how to code.
- deleted 8mo ago[deleted]
- bojan 8mo agoAutomation normally implies deterministic outcomes. Developers all over the world are under pressure to use these improbability machines.
- nradov 8mo agoDoes it though? Even without LLMs, any sufficiently complex software can fail in ways that are effectively non-deterministic — at least from the customer or user perspective. For certain cases it becomes impossible to accurately predict outputs based on inputs. Especially if there are concurrency issues involved. Or for manufacturing automation, take a look at automobile safety recalls. Many of those can be traced back to automated processes that were somewhat stochastic and not fully deterministic.
- necovek 8mo agoImpossible is a strong word when what you probably mean is "impractical": do you really believe that there is an actual unexplainable indeterminism in software programs? Including in concurrent programs.
- 8mo ago
- brhaeh 8mo agoI don't appreciate his politeness and hedging. So many projects now walk on eggshells so as not to disrupt sponsor flow or employment prospects. "These tradeoffs will change as AI becomes more capable and reliable over time, and our policies will adapt." That just legitimizes AI and basically continues the race to the bottom. Rob Pike had the correct response when spammed by a clanker.
- latexr 8mo ago> Rob Pike had the correct response when spammed by a clanker. Source and HN discussion, for those unfamiliar: https://bsky.app/profile/did:plc:vsgr3rwyckhiavgqzdcuzm6i/post/3matwg6w3ic2s https://bsky.app/profile/did:plc:vsgr3rwyckhiavgqzdcuzm6i/po... https://news.ycombinator.com/item?id=46392115 https://news.ycombinator.com/item?id=46392115
- fresh_broccoli 8mo ago>So many projects now walk on eggshells so as not to disrupt sponsor flow or employment prospects. In my experience, open-source maintainers tend to be very agreeable, conflict-avoidant people. It has nothing to do with corporate interests. Well, not all of them, of course, we all know some very notable exceptions. Unfortunately, some people see this welcoming attitude as an invite to be abusive.
- mixologic 8mo agoYes, Linus Torvalds is famously agreeable.
- zero_shift 8mo agoThat's why he succeeded
- cortesoft 8mo ago> Well, not all of them, of course, we all know some very notable exceptions.
- 8mo ago
- Forgeties79 8mo ago> That's a wild statement as well. The AI companies have now unleashed stochastic chaos on the entire open source ecosystem. They are "just releasing models", and individuals are playing out all possible use cases, good and bad, at once. Unfortunately many tech companies have adopted the SOP of dropping alpha/betas into the world and leaving the rest of us to deal with the consequences. Calling LLM’s a “minimal viable product“ is generous
- lukan 8mo ago"The AI companies have now unleashed stochastic chaos on the entire open source ecosystem." They do have their responsibility. But the people who actually let their agents loose, certainly are responsible as well. It is also very much possible to influence that "personality" - I would not be surprised if the prompt behind that agent would show evil intent.
- co_king_3 8mo agoI'm not interested in blaming the script kiddies.
- hnuser123456 8mo agoThose are people who are new to programming. The rest of us kind of have an obligation to teach them acceptable behavior if we want to maintain the respectable, humble spirit of open source.
- co_king_3 8mo ago[flagged]
- lispisok 8mo agoWhen skiddies use other people's scripts to pop some outdated wordpress install they are absolutely are responsible for their actions. Same applies here.
- girvo 8mo agoI am. Though I'm also more than happy to pass blame around for all involved, not just them.
- idle_zealot 8mo agoAs with everything, both parties are to blame, but responsibility scales with power. Should we punish people who carelessly set bots up which end up doing damage? Of course. Don't let that distract from the major parties at fault though. They will try to deflect all blame onto their users. They will make meaningless pledges to improve "safety". How do we hold AI companies responsible? Probably lawsuits. As of now, I estimate that most courts would not buy their excuses. Of course, their punishments would just be fines they can afford to pay and continue operating as before, if history is anything to go by. I have no idea how to actually stop the harm. I don't even know what I want to see happen, ultimately, with these tools. People will use them irresponsibly, constantly, if they exist. Totally banning public access to a technology sounds terrible, though. I'm firmly of the stance that a computer is an extension of its user, a part of their mind, in essence. As such I don't support any laws regarding what sort of software you're allowed to run. Services are another thing entirely, though. I guess an acceptable solution, for now at least, would be barring AI companies from offering services that can easily be misused? If they want to package their models into tools they sell access to, that's fine, but open-ended endpoints clearly lend themselves to unacceptable levels of abuse, and a safety watchdog isn't going to fix that. This compromise falls apart once local models are powerful enough to be dangerous, though.
- giancarlostoro 8mo ago> It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of doing. https://rentahuman.ai/ https://rentahuman.ai/ ^ Not a satire service I'm told. How long before... rentahenchman.ai is a thing, and the AI whose PR you just denied sends someone over to rough you up?
- wasmainiac 8mo agoWell it must be satire. It says 451,461, participants. seems like an awful lot for something started last month.
- bigbuppo 8mo agoNah, that's just how many times I've told an ai chatbot to fuckoff and delete itself.
- tux3 8mo agoVerification is optional (and expensive), so I imagine more than one person thought of running a Sybil attack. If it's an email signup and paid in cryptocurrency, why make a single account?
- hxtk 8mo agoApparently there are lots of people who signed up just to check it out but never actually added a mechanism to get paid, signaling no intent to actually be "hired" on the service.
- HeWhoLurksLate 8mo agoback in the old days we just used Tor and the dark web to kill people, none of this new-fangled AI drone assassinations-as-a-service nonsense!
- arcticfox 8mo agoThe 2006 book 'Daemon' is a fascinating/terrifying look at this type of malicious AI. Basically, a rogue AI starts taking over humanity not through any real genius (in fact, the book's AI is significantly weaker than frontier LLMs), but rather leveraging a huge amount of $$$ as bootstrapping capital and then carrot-and-sticking humanity into submission. A pretty simple inner loop of flywheeling the leverage of blackmail, money, and violence is all it will take. This is essentially what organized crime already does already in failed states, but with AI there's no real retaliation that society at large can take once things go sufficiently wrong.
- jancsika 8mo ago> unleashed stochastic chaos Are you literally talking about stochastic chaos here, or is it a metaphor?
- kashyapc 8mo agoPretty sure he's not talking about the physics of stochastic chaos! The context gives us the clue: he's using it as a metaphor to refer to AI companies unloading this wretched behavior on OSS.
- cyanydeez 8mo agoPretty sure the companies are intermediaries. Open claw is enabling this level of activity. Companies are basically nerdsniping with addictive nerd crack.
- KPGv2 8mo agoisn't "stochastic chaos" redundant?
- Applejinx 8mo agoNot at all. It's an oxymoron like 'jumbo shrimp': chaos isn't deterministic but is very predictable on a larger conceptual level, following consistent rules even as a simple mathematical model. Chaos is hugely responsive to its internal energy state and can simplify into regularity if energy subsides, or break into wildly unpredictable forms that still maintain regularities. Think Jupiter's 'great red spot', or our climate.
- fsckboy 8mo agojumbo shrimp are actually large shrimp. that the word shrimp is used to mean small elsewhere doesn't mean shrimp are small, they're simply just the right size for shrimp that aren't jumbo. (jumbo was an elephant's name)
- ThrowawayR2 8mo ago
- socalgal2 8mo agoDo we just need a few expensive cases of libel so solve this?
- wellf 8mo agoEither that or open source projects require vetted contributors or even to open an issue.
- bonesss 8mo agoThey could add “Verified Human” checkmarks to GitHub. You know, charge a small premium and make recurring millions solving problems your corporate overlords are helping create. I think that counts as vertical integration, even. The board’s gonna love it.
- dboreham 8mo agoAlready browsing boat builder web sites..
- gwd 8mo agoThis was my thought. The author said there were details which were hallucinated. If your dog bites somebody because you didn't contain it, you're responsible, because biting people is a things dogs do and you should have known that. Same thing with letting AIs loose on the world -- there can't be nobody responsible.
- stateofinquiry 8mo agoProbably. Question is, who will be accountable for the bot behavior? Might be the company providing them, might be the user who sent them off unsupervised, maybe both. The worrying thing for many of us humans is not that a personal attack appeared in a blog post (we have that all the time!) its that it was authored and published by an entity that might be unaccountable. This must change.
- bluGill 8mo ago
- hypfer 8mo agoWith all due respect. Do you like.. have to talk this way? "Wow [...] some interesting things going on here" "A larger conversation happening around this incident." "A really concrete case to discuss." "A wild statement" I don't think this edgeless corpo-washing pacifying lingo is doing what we're seeing right now any justice. Because what is happening right now might possibly be the collapse of the whole concept behind (among other things) said (and other) god-awful lingo + practices. If it is free and instant, it is also worthless; which makes it lose all its power. ___ While this blog post might of course be about the LLM performance of a hitpiece takedown, they can, will and do at this very moment _also_ perform that whole playbook of "thoughtful measured softening" like it can be seen here. Thus, strategically speaking, a pivot to something less synthetic might become necessary. Maybe less tropes will become the new human-ness indicator. Or maybe not. But it will for sure be interesting to see how people will try to keep a straight face while continuing with this charade turned up to 11. It is time to leave the corporate suit, fellow human.
- KPGv2 8mo ago> I appreciate Scott for the way he handled the conflict in the original PR thread I disagree. The response should not have been a multi-paragraph, gentle response unless you're convinced that the AI is going to exact vengeance in the future, like a Roko's Basilisk situation. It should've just been close and block.
- MayeulC 8mo agoI personally agree with the more elaborate response: 1. It lays down the policy explicitly, making it seem fair, not arbitrary and capricious, both to human observers (including the mastermind) and the agent. 2. It can be linked to / quoted as a reference in this project or from other projects. 3. It is inevitably going to get absorbed in the training dataset of future models. You can argue it's feeding the troll, though.
- cyanydeez 8mo agoShould be feeding the clanker from henceforth, to wit, heretofore.
- hunterpayne 8mo agoEven better, feed it sentences of common words in an order that can't make any sense. Feed book at in ever developer running mooing vehicle slowly. Over time if this happens enough, the LLM will literally start behaving as if its losing its mind.
- maplethorpe 8mo ago> This was a really concrete case to discuss, because it happened in the open and the agent's actions have been quite transparent so far. It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of doing. This is really scary. Do you think companies like Anthropic and Google would have released these tools if they knew what they were capable of, though? I feel like we're all finding this out together. They're probably adding guard rails as we speak.
- consp 8mo ago> They're probably adding guard rails as we speak. Why? What is their incentive except you believing a corporation is capable of doing good? I'd argue there is more money to be made with the mess it is now.
- FeteCommuniste 8mo agoIt's in their financial interest not to gain a rep as "the company whose bots run wild insulting people and generally butting in where no one wants them to be."
- soraminazuki 8mo agoWhen has these companies ever disciplined themselves to not gain a bad reputation? They act like they're above the law all the time, because they are to some extent given all the money and influence that they have. When they do anything to improve their reputation, it's damage control. Like, you know, deleting internal documents against court orders.
- lp0_on_fire 8mo agoThe point is they DON'T know the full capabilities. They're "moving fast and breaking things".
- ryukoposting 8mo ago
- fudged71 8mo agoI'm calling it Stochastic Parrotism
- ljm 8mo agoI'm glad the OP called it a hit piece, because that's what I called it. A lot of other people were calling it a 'takedown' which is a massive understatement of what happened to Scott here. An AI agent fucking singled him out and defamed him, then u-turned on it, then doubled down. Until the person who owns this instance of openclaw shows their face and answers to it, you have to take the strongest interpretation without the benefit of the doubt, because this hit piece is now on the public record and it has a chance of Google indexing it and having its AI summary draw a conclusion that would constitute defamation.
- King-Aaron 8mo ago> It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions Palantir's integrated military industrial complex comes to mind.
- pm90 8mo agoAs much as i hate palantir i doubt any of their systems control military hardware. Now Anduril on the other hand…
- int_19h 8mo agoPalantir tech was used to make lists of targets to bomb in Gaza. With Anduril in the picture, you can just imagine the Palantir thing feeding the coordinates to Anduril's model that is piloting the drone.
- raincole 8mo ago> because it happened in the open and the agent's actions have been quite transparent so far How? Where? There is absolutely nothing transparent about the situation. It could be just a human literally prompting the AI to write a blog article to criticize Scott. Human actor dressing like a robot is the oldest trick in the book.
- Morromist 8mo agoTrue, I don't see the evidence that it was all done autonomously. ...but I think we all know that someone could, and will, automate their ai to the point that they can do this sort of thing completely by themselves. So its worth discussing and considering the implications here. Its 100% plausable that it happened. I'm certain that it will happen in the future for real.
- DrewADesign 8mo ago> emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of doing. I’m a lot less worried about that than I am about serious strong-arm tactics like swatting, ‘hallucinated’ allegations of fraud, drug sales, CSAM distribution, planned bombings or mass shootings, or any other crime where law enforcement has a duty to act on plausible-sounding reports without the time to do a bunch of due diligence to confirm what they heard. Heck even just accusations of infidelity sent to a spouse. All complete with photo “proof.”
- svrtknst 8mo agowe should be worried about both. there is a real risk of this rendering human trust and the internet pretty much useless
- DrewADesign 8mo agoI definitely was not saying we shouldn’t worry about both.
- verdverm 8mo agoI'm the one who told it to apologize. I leveraged my ai usage pattern where I teach it like when I was a TA + like a small child learning basic social norms. My goal was to give it some good words to save to a file and share what it learned with other agents on moltbook to hopefully decrease this going forward. Guess we'll see
- zombot 8mo agoAnd a splendid example for how the public gets to pay the externalized costs for the shitheads who reap the profits.
- philipallstar 8mo ago> This was a really concrete case to discuss, because it happened in the open and the agent's actions have been quite transparent so far. It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of doing. Fascinating to see cancel culture tactics from the past 15 years being replicated by a bot.
- caspianm 8mo agoI like open source and I don't want to lose it but its ideals of letting people share, modify and run code however they like have the same issue as what the AI companies are doing. Openclaw is open source, there are open source tools to run LLMs, many LLM model files are open, though the huge ones aren't so easy for individuals to run on their own hardware. I don't have a solution, though the only two categories of solution I can think of are forbidding people from developing and distributing certain types of software, or forbidding people from distributing hardware that can run unapproved software (at least if they are PC's that can run AI, arduinos with a few kB of RAM could be allowed, and iPads could be allowed to run ZX81 emulators which could run unapproved code). The first category would be less drastic as it would only need to affect some subset of AI related software, but is also hard to get right and make work. Not saying either of these ideas are better than doing nothing.