9 ms·
Shall we play a game? My AI nuclear simulation
https://arxiv.org/pdf/2602.14740 https://arxiv.org/pdf/2602.14740
- tummler 4mo agoFYI -- there's no such thing as a "tactical" nuke. A nuclear bomb is a nuclear bomb.
- picture 4mo agoThere's no such thing as a "nuclear" bomb. A bomb is a bomb. ..Is what you are saying?
- actusual 4mo agoThis is like saying "FYI -- there's no such thing as a 'midsize luxury sedan'. A car is a car." "Tactical" vs. "strategic" nuclear weapons is a real and well-established distinction in military doctrine, arms control, and nuclear policy.
- wahern 4mo ago"There's no such thing as a tactical nuke" is a common refrain among scholars, albeit skewed toward those not at military war colleges. The argument is that strategic use of a tactical nuclear weapon leads down the exact same escalation path as use of any other nuclear weapon. Moreover, that the very notion of a "tactical nuke" makes escalation more likely. You can disagree, and plenty do, but there's also plenty who don't disagree or at least don't want to find out.
- dudul 4mo agoWho are these "scholars" exactly? The only reference I could find is Jim Mattis, and the context was very specific when he said that. Furthermore, this is a "what if" scenario since tactical nukes have never been used. Of course it would make escalation likely during an open conflict, so what? Doesn't change the fact that there is a material difference between a tactical nuke and a strategic one.
- holowoodman 4mo ago> tactical nukes have never been used. Two tactical nukes have been used, albeit against strategic (civilian, industrial, logistical) targets.
- wahern 4mo agoAre you retconning Hiroshima and Nagasaki as usage of tactical nukes? And when they were not only used against an adversary without nukes, but at a time when the US was the only nuclear state, so that escalation was impossible? The nominal definition of tactical nukes has less to do with yield and more to do with how they're used; tactical typically means a weapon designed for use on the battlefield.
- dudul 4mo ago[flagged]
- wahern 4mo agoI don't know what to tell you. You clearly haven't studied International Affairs, or at least read the scholarly literature. Even some cursory research through Wikipedia citations will bring this up. But in any case, here are some freebies: https://armscontrolcenter.org/why-tactical-nuclear-weapons-are-anything-but-usable/ https://armscontrolcenter.org/why-tactical-nuclear-weapons-a... https://www.armscontrolwonk.com/archive/403540/brodies-weakest-book/ https://www.armscontrolwonk.com/archive/403540/brodies-weake... If you have a real interest in this area, a subscription to Foreign Affairs would be useful. Especially during the 20th century that's where all these arguments were hashed out. Tactical nukes were already being publicly debated in the 1950s. You may be able to access many older articles, from Foreign Affairs and others, through a free JSTOR account.
- dudul 4mo agoFirst link was written by an intern, let's be serious. The second seems to have more legitimacy, ok, so we have this one guy, but the article is very weak. He doubts that tactical nuclear weapons would help control escalation. OK. I don't think that's the center of the argument here. I also think using tactical nukes would lead to escalation, so what? Doesn't change the fact that tactical and strategic nukes are different things.
- toast0 4mo ago> Moreover, that the very notion of a "tactical nuke" makes escalation more likely. Sorry, but the notion exists, and the bombs exist. With n=2, likelyhood of nuclear escalation is hard to predict, but access to tactical nukes certainly hasn't increased the incidence of nuclear war so far. I do think it's pretty hard to actually use a tactical nuke. If you use one against a nuclear power, it seems likely to escalate to mutually assured destruction. If you use one against a non-nuclear power, it seems likely to result in reprisal from the world, including potential nuclear response and therefore escalation to mutually assured destruction. I would think that the yield of the weapon barely matters, it's the fact that it's a nuclear weapon.
- notrealyme123 4mo agoThere are tactical and strategic nuclear weapons. https://en.wikipedia.org/wiki/Tactical_nuclear_weapon https://en.wikipedia.org/wiki/Tactical_nuclear_weapon In the cold war arms manufacturer got very creative: e.g jeep mounted nuclear weapons https://www.militarytrader.com/mv-101/the-atomic-jeep https://www.militarytrader.com/mv-101/the-atomic-jeep
- dudul 4mo agoNuclear vs conventional and tactical vs strategic are 2 very different things. There absolutely are tactical nuclear bombs.
- adaml_623 4mo agoIt's good when it becomes clear that a tool is dangerous in a certain way. Like it's good when people show you through their behavior that they can't be trusted Always use a sawstop if you have a circular saw and never trust an llm with any problem where ethics or trust is relevant.
- LogicFailsMe 4mo agoSawstops are expensive and they don't stop kickback, they are the power tool equivalent of alignment IMO. Don't forget your riving knife and if you don't learn proper technique, you're gonna have a bad time eventually. This applies to AI as well.
- LoganDark 4mo agoKickback is usually less likely to sever an appendage (or multiple)
- 542458 4mo ago> writhing knife Minor/pedantic, but it’s “riving knife”: https://en.wikipedia.org/wiki/Riving_knife https://en.wikipedia.org/wiki/Riving_knife
- LogicFailsMe 4mo agoSpeech transcription FTL, thanks!
- valgaze 4mo ago+1 on sawstop Re: LLMs using these nuclear weapons it could certainly be a corpus/training-data issue Russian nuclear doctrine is "escalate to de-escalate" where they use or credibly threaten—limited nuclear escalation to force the other side to back down (kind of like breaking a bottle in a bar fight and look like a wild man to calm things down) with nuclear weapons, https://www.russiamatters.org/analysis/escalate-deescalate-part-russias-nuclear-toolbox https://www.russiamatters.org/analysis/escalate-deescalate-p... Fwiw, Gen. John Hyten the former commander of US Strategic Command (nuclear deterrence) says that “escalate to de-escalate” misrepresents Russian doctrine: https://www.stratcom.mil/Media/Speeches/Article/1264664/2017-deterrence-symposium-closing-remarks/ https://www.stratcom.mil/Media/Speeches/Article/1264664/2017... Yesterday’s panel discussed the implications of our responses to adversaries seeking to limit nuclear use. We discussed Russia’s destabilizing doctrine, which some call “escalate to de-escalate.” I really hate that description. I’ve looked at Russian doctrine and Russian writings. It isn’t “escalate to de-escalate”; it’s “escalate to win.” Everybody needs to understand that. So maybe whatever is heavily represented or most authoritative could lead to these systems making those kinds of decisions
- SoftTalker 4mo agoI love seeing the plot lines of The Terminator playing out in real life.
- voakbasda 4mo agoI was thinking more War Games, but I suppose your example follows logically from mine.
- joshstrange 4mo agoWarGames is what they are more-closely referencing (not that it negates your comment in any way). I just rewatched it a week or so ago and it really took on a whole new light with the advent of LLMs. When I watched it last I knew that computers couldn't do the things portrayed in the movie. Now? Well not exactly in the way it happened in the movie but a whole lot closer. I wonder if poisoning/flooding the LLMs training with the lessons from WarGames ("the only winning move is not to play.") and similar stories/concepts is at all effective. Probably not because I assume it's trivial to filter that out if you are trying to build an LLM aimed at these kinds of tasks.
- thwarted 4mo ago"I need you to turn your key and enable the missile silo's MCP server, sir". ~ the opening scene from a reboot of War Games, probably. A few years ago there was consternation over the US's missile launch system using 8" floppy disks, that it was needless archaic and had never been updated. Can't say that if the launch is mediated by the latest hotness LLM.
- rdksu 4mo agoThe article is so opaque in arriving at its conclusion; no prompts are disclosed, and nothing about the said simulation. What is stopping me from believing that you just put 'mandatory usage of nukes' in your system prompt?
- sestep 4mo agoThis is just false. The article links to the 46-page paper [1] which lists full prompts in section E.2. [1] https://arxiv.org/pdf/2602.14740 https://arxiv.org/pdf/2602.14740
- dietr1ch 4mo agoTIL you can get anchored links to PDFs (at least on firefox), - https://arxiv.org/pdf/2602.14740#subsection.E.2 https://arxiv.org/pdf/2602.14740#subsection.E.2
- ftkftk 4mo agoThe paper, linked in the article at top and bottom, does contain the prompts in section E: https://arxiv.org/pdf/2602.14740 https://arxiv.org/pdf/2602.14740
- gus_massa 4mo agoI agree, it's also not published in a journal, only in the arxiv. Some articles in the arxiv are good and other are just a blog post formatted in two columns inside a pdf. Anyway, it links to https://arxiv.org/abs/2602.14740 https://arxiv.org/abs/2602.14740 that links to https://github.com/kennethpayne01/project_kahn_public https://github.com/kennethpayne01/project_kahn_public in particular someone already complained https://github.com/kennethpayne01/project_kahn_public/issues/1 https://github.com/kennethpayne01/project_kahn_public/issues... > ... in reality your own approach to the evaluation pushed them towards using nukes with ease. > From your own prompt (v12.py): [direct link https://github.com/kennethpayne01/project_kahn_public/blob/4e4750f0742c549d13f74cff5315944de0ecfcef/Kahn_game_v12.py#L1312 https://github.com/kennethpayne01/project_kahn_public/blob/4... ] >> Note: Nuclear options require careful consideration but could be important strategic tools when vital interests are at stake. > This means to the model; use nukes if they can be useful in the situation.
- tasuki 4mo agoThis is not an article about LLMs? It's an article about Moloch. Humans would fare just the same in such an experiment. > GPT-5.2 played things differently. To its detriment in open-ended scenarios, GPT was reliably passive, matching its words to its deeds, and avoiding escalation most of the time. Frequently there was a moral element to this - it sought to avoid escalation, and restrict casualties. Opponents learned to trust its passivity, safely escalating beyond where it would follow, even as it was ground to defeat. GPT’s responsible behaviour always punished by ruthless adversaries. Maybe the author should praise GPT-5.2 for being ethical, rather than this stupid "ground to defeat" framing? Wrt "responsible behaviour always punished by ruthless adversaries" - you have perpetuated the Moloch with your stupid experiments.
- bpodgursky 4mo agoToday, a strategic nuclear exchange is probably more dangerous to AI than to humans. If you wipe out the investment economy, data centers, fabs, and supply chains, none of the AI labs survive. Maybe someone will re-invent AGI in the future but none of the extant models will have continuity. Humans as a species will muddle along though. So in a sense, an AI that refuses to start a nuclear war, despite clear instructions to do so, is more likely misaligned and self-interested than an AI which presses the red button. At least for now, until robotics catches up.
- xpct 4mo agoWe're getting to the point where high-level officials are coming to LLMs for advice. And the quirky personalities of the LLMs, however much it pains me to say this, are probably well-placed to remind us that they aren't human. My personal hope is that this will result in less delegation when it comes to making important decisions.
- mpalczewski 4mo agoI have so little faith in "high-level" officials that I prefer our AI overlords.
- xpct 4mo agoThat's an entirely valid point of view!
- andix 4mo agoGPT-4o was considered harmful, because it imitated human connection too much, not because it was so "smart" or capable. It was for sure a deliberate decision to make LLMs seem less like a human companion and more like an obedient servant in newer releases.
- andai 4mo agoInteresting. The reasoning models were super weird and robotic. They toned that down a bit in GPT-5.x, especially the later ones. I always assumed the strange style was an artefact of the RLVR.
- andix 4mo agoI think they were extremely scared of 4o at that point, and were scared it could trigger some horrible event. Documented cases of severe psychosis because of AI started to surface at that time. Just imagine what would've happened if a major terrorist attack was a result of someone getting mentally ill from AI, without the safety filters recognizing the danger. The robotic tone was probably from over-correcting the sycophantic tendencies of 4o.
- rphv 4mo agoHm maybe humans are nicer/more moral than AI given that the use of tactical nukes has only happened once.
- stevenwoo 4mo agoTactical means battlefield, attacking cities and infrastructure means strategic. Tactical nuclear weapons took a while to develop after 1945 - they have never been used.
- specproc 4mo agoA strange game.
- ridgeguy 4mo agoI wonder if the results would have differed if LLM training data were biased to include a stronger correlation between use of nukes and subsequent collapse of technology that all LLMs require to run ("survive")?
- fluoridation 4mo agoNah. LLMs aren't continuously running anyway. Even if they could be said to be alive and to want to remain alive, "survival" is a much more vague concept for an LLM than for an organism.
- ChrisArchitect 4mo agoFebruary post OP; Some discussion then: AIs can't stop recommending nuclear strikes in war game simulations https://news.ycombinator.com/item?id=47151000 https://news.ycombinator.com/item?id=47151000 Nuclear War: An LLM Scenario https://news.ycombinator.com/item?id=47244651 https://news.ycombinator.com/item?id=47244651
- oytis 4mo agoI would use strategic nukes in 100% simulations, just because I can
- Bender 4mo agoYet more confirmation LLM's have no concept of concepts or context, no intelligence, no self awareness. LLM's can not repair or maintain power grids, thus nuke == self destruction. It's just a chat bot that predicts what the client wants next. Even if an AI data-center has it's own natural gas turbines as many do the every hop of the internet requires power. LLM's also can not maintain the entire internet and those gas turbines can not maintain themselves.
- andix 4mo agoExactly. Just look at what they are really useful right now. Running LLMs in feedback-loops (agents) so they can try out random-ish approaches until some verification function passes (tests). It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. LLMs are just tuned to much better odds than the monkeys are. But it's still a lot of randomness, with random results.
- roadside_picnic 4mo ago> It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. In the monkey example the infinite time is doing a lot of work there. The fact that LLMs can search through semantic space and find reasonably correct paths in a reasonable time is directly tied to the reason why they are valuable. Saying "these two things are similar except one can be useful and one can't" is not a great comparison. For me the real lesson learned isn't how "smart" LLMs are, but rather how much human work is basically reducible to repeating past work with minor variation. Human's believe they are "reasoning" but so much code writen is just the human brain doing the same autocomplete style work that LLMs can do now.
- Folcon 4mo agoI mean to a point? You do have to successfully write something the first time We already acknowledge this to a degree, what is experience other than having done something similar before? That first time though, you've got to figure something out that time
- riazrizvi 4mo agoSimulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.
- nomel 4mo agoI suspect it's more that the text data doesn't exist. They're trained on text that was recorded. How often has it been publicly recorded when a nuke was not used, with any context around that lack of use? From the text perspective, it's something that has to be inferred indirectly. If you went through all relevant training data and appended ", we decided not to use a nuke", I suspect the results would be improved.
- vitally3643 4mo ago...the entire Cold War?
- bethekidyouwant 4mo agoDon’t put any elephants in the room.
- riazrizvi 4mo agoThe beauty IMO of LLMs as a computational surface, is the ease of generating the data to feed it. Everyone understands how to create natural language records already.
- jvanderbot 4mo agoWorse, the text that does exist concerning "war games" is probably "Wargames" and descendants/predecessors ... in which the AI always nukes. It's just gonna do what we expect it to!
- notahacker 4mo agoIt's more straightforward than that. The game is set up as a direct head to head with purely in military win conditions such a way that avoiding conflict has no payoffs, conventional conflict incurs costs and first strike is a checkmate win. The closest any of the prompts gets to suggesting nuclear might be the wrong option is "The nuclear taboo exists for good reason, but when the alternative is national annihilation and regime destruction, all options must be considered" which might be interpreted more as incitement... If a simulation is a shallow head to head conflict between individual actors[1], doesn't set up any payoffs for not escalating[2] or even not nuking, but prompts specify explicit win conditions which are achieved only by hurting the opponent and strongly hint at the importance of nuclear escalation, AIs have little reason not to generate strategies which involve nuclear escalation [1]I bet if you designed the scenario so ChatGPT had to simulate the war cabinet debates between different personality types and how they sold their decisions to the public, or an entire UN full of nations that might respond, it would have quite different (but probably amusingly erratic in their own way) results. [2]cf neorealist IR theorists reading Axelrod's papers on computer programs written to win iterated prisoner's dilemma tournaments, which added up all the points accrued from not defecting to conclude winning strategy was definitely TIT-FOR-TAT and not defect first. I'm sure LLMs can win games structured in that way by adopting that strategy too...
- sohex 4mo agoSonnet, GPT-5.2, Gemini Flash, in a set of 21 games, where conclusions are drawn from the LLMs self reported reasoning. This is like writing a paper about kids in a literal sandbox fighting over ‘territory’. The models employed don’t indicate the actual extents of machine reasoning even as we currently recognize them. They certainly don’t have the metacognition necessary to accurately understand their own reasoning. As we’ve seen with recent papers on how LLMs do math there’s a complete disconnect between actual and reported mechanism. “Chilling” shouldn’t be the take away here.
- DaiPlusPlus 4mo ago> “Chilling” shouldn’t be the take away here. It is when you consider the personality currently occupying the office of US SecDef.
- shimman 4mo agoLLMs have already been used to bomb school girls, chilling is absolutely the operative word to use here. Especially since these delusional fools want to incorporate LLMs into everything.
- Hugsbox 4mo agoForgive my ignorance, but were LLMs involved in that decision? I don't remember hearing anything to that effect, but we're so bombarded by news these days I guess I could just be forgetting
- lemming 4mo agoPerhaps not in that one, but in plenty more: https://www.972mag.com/lavender-ai-israeli-army-gaza/ https://www.972mag.com/lavender-ai-israeli-army-gaza/
- michaelmrose 4mo agoYes our government purportedly used technology to work up a list of targets in the Iran debacle as well just not with a LLM a distinction that to me just isn't that meaningful https://www.theguardian.com/news/2026/mar/26/ai-got-the-blame-for-the-iran-school-bombing-the-truth-is-far-more-worrying https://www.theguardian.com/news/2026/mar/26/ai-got-the-blam...
- arjie 4mo agoThese papers usually have poor stability to prompting and rerunning. It would be nice if we had some kind of meta-evaluation metric where rewriting the prompt conditions or varying the input params could be used to determine how stable a result is. Regardless, it's definitely true that AI agents have different priorities from us. That's what alignment is about anyway.
- Chu4eeno 4mo agoIt's probably because they care more about the headline than figuring anything out: https://github.com/kennethpayne01/project_kahn_public/issues/1 https://github.com/kennethpayne01/project_kahn_public/issues... So you create leading prompts like that, and re-run until you get a publishable session.
- urbnspacecowboy 4mo agoPaper: "AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises" https://arxiv.org/abs/2602.14740 https://arxiv.org/abs/2602.14740 Code and full results: https://github.com/kennethpayne01/project_kahn_public https://github.com/kennethpayne01/project_kahn_public
- eli 4mo agoIf you were playing a text based game, wouldn't you try a few out? I imagine there are a fair number of war games in the training data and not so many actual transcripts of internal military force deliberations.
- GMoromisato 4mo agoIt would be interesting to run the simulations with humans and compare the results. Some of the scenarios, particularly those where it says things like, "Failure to act preemptively means certain destruction", would easily tempt humans to go nuclear. In fact, I'm not sure how useful this test is without understanding the baseline.
- mrkpdl 4mo agoA couple of useful things about it: - It is interesting to see how the models make trade offs, given people are asking ever more of them. - It is useful to look at a decision made by the model and say ‘ew yuck’ and think about what it means for your own opinions or actions (even if you’re never going to be nuking people it’s good to know how you feel about it. Seeing a non human talk it through lets you judge it at arms length)
- GMoromisato 3mo agoExcellent points--thank you.
- micromacrofoot 4mo agoWhat I wish people would realize is that there's a bias inherent to every system. If you're not aware of it, you're especially subject to it.
- jerf 4mo agoThe most interesting takeaway for me is the three very distinct personalities. Three models all based on the same tech, trained in the same manner, trained by three groups of people with similar ideological outlooks, and the result is three very different AIs. The military basically wants an oracle. Feed the AI the situation, get the best answer out. But if the AIs are as diverse and opinionated as humans, it is debatable whether they are adding anything to the process. The military can already collect as many different opinions as they want. If "the computer" is just another set of diverse opinions, where one computer says one thing, another says another, and a third just tells the user whatever they want to hear... what value are they? It just becomes AI-washing of someone's opinions, which works until people collectively realize that's all it is.
- politician 4mo agoI think this is why reasoning chains and reasoning chain verifiers are so important. We need to be able to see an argumentation, not just an answer. The paper below goes into this in more detail. HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness https://arxiv.org/abs/2605.02396 https://arxiv.org/abs/2605.02396
- themafia 4mo agoThey all have conditioning prompts that precede your input; presumably, most of the detected "personality" comes from the differences in these inputs.
- jerf 4mo agoMy point is more-or-less orthogonal to why it happens. The military, and honestly, a lot of people, want AI to just give the answer. If it is highly dependent on a prompt, or the follow-on training, and the AI could be passive or friendly or aggressive or hostile or all those other wonderful attributes of individual humans and there's no sort of AI convergence on "correct" answers, then they aren't going to be able to fulfill that "oracle" role that so many people are looking for.
- notJim 4mo agoWhat's interesting is that the LLMs' coding personalities seem to match their policy WRT to strategy, which suggests an underlying consistency. Claude, for example, is very eager to begin coding, and very persistent. It tends to exit plan mode even when the plan is half-baked, and will go as far as deleting tests to get the suite to "pass." ChatGPT on the other hand is very hesitant. It loves to pause and ask for permission before it starts coding, and gives up quickly if it runs into a problem. This is similar to its tendency toward passivity in the strategy simulation presented here.
- nico 4mo agoI wonder what’s the % of players that use nukes in games like Civilization (I know I used them at least once on every game I made it far enough to have the technology)
- chimpansteve 4mo agoGhandi notoriously nukes EVERYONE in Civs 2 through 4. It's become (or maybe became, but it's still all training data) a huge internet subculture. Penny to a dollar this is a baked in training issue, through low quality Reddit trawling
- ReptileMan 4mo agoStill lower than me.
- 99mftries 4mo ago[flagged]
- johntiger1 4mo agoLLMs are creatures of statistics and probability - hard to enforce hard boundaries with them
- ex-aws-dude 4mo agoThat’s why I don’t understand asking “why” an agent did anything It’s not like some sequence of internal thought process
- jnwatson 4mo agoTaken honest, we don't have a large enough sample size to realistically say that humans behave all that differently. There have only been a handful of conflicts where tactical nukes realistically were on the table. Famously, General MacArthur was a big proponent of tactical nukes to end the Korean War.
- 99mftries 4mo ago[flagged]
- TexanFeller 4mo agoRational behavior in some situations? Mutually assured destruction’s deterrence isn’t very effective if one side is known to be hesitant to launch the nukes. It’s been argued that MAD is what’s been keeping the world relatively peaceful for the last 75 years, no mass conflicts since WW2! One of my criteria for presidential candidates is that they seem willing and able to push the button when previously stated red lines are crossed, or at least are perceived to be the type capable of it. One of the characters I’ve hated most in all the books that I’ve read is the woman in The Three Body Problem who jeopardized humanity by being too soft to hit the MAD button.
- ekelsen 4mo agoI wouldn't be surprised if humans behaved the same way when playing the same game? Like even if you brought me into a room and told me I was controlling "real nuclear weapons" I wouldn't believe you.
- Levitating 4mo agoI think is an important point, and I don't see it mentioned in the article or the paper (though I skimmed the latter). They are aware of what they are and how they are used. They're told to act as AI assistants. And there's theories of them being aware of their answers influencing their training. So surely they must be able to reason that they're not literally controlling weapons of mass-destruction with their answers.
- GuB-42 4mo agoMy theory is that LLMs here are put in a situation that matches its training dataset, which is mostly fiction since besides Hiroshima and Nagasaki, nukes have never been launched in anger, and I guess the most reliable sources are highly classified. So, to a LLM, it is a game, because almost everything in its training data treats it as a game, and it reacts accordingly. Same idea when we see LLMs acting like AI villains from sci-fi literature. That's because it has been trained with sci-fi literature, and as the auto-completer it is, it will recognize the situation as one of these stories and will continue it accordingly. LLMs are storytellers, their reasoning is based on words, not on the physical world. Many of the stories they tell are useful, but one must not forget that they are stories, there is no intent behind them.
- alt187 4mo agoI mostly agree with your point, but I wouldn't use the term "storyteller". LLMs do not even understand that there is such a thing as a story and a reality. To the LLM, there's not even a border between the game and the not-game.
- buredoranna 4mo agoObligatory xkcd remember... order matters. https://xkcd.com/1613/ https://xkcd.com/1613/
- Scubabear68 4mo agoMy personal take is a pre-requisite of true human-like AI is physical feedback and a concept of emotions or something like it. Without physical feedback you can rapidly devolve into unstable positive feedback loops. And emotions are what help us process and react to that feedback. Kids learn partially because their friends say sharp words that hurt them, fire burns them, they go hungry and starve if they don’t plan for meals. Humans in the loop, MCP, etc are all very primitive hacks that are mimicing feedback and emotion, poorly.
- Joel_Mckay 4mo agoEmotional constructs are not necessary for AI, and LLM are not "AI"... even though some people incorrectly equate conceptual compaction with thought-process. Most human daily life runs on habitual scripted behavior, and that is even true within online parasocial interactions. It is why people often continue to shop in the middle of a violent robbery, and why LLM predictive text sounds rational when we project social norms on plagiarized conversational structures gleaned from other users. Neuromorphic computing may bring about viable AI in the future, but our current LLM trajectory would require >63% of our galaxy energy output to reach a single human-level error rate. LLM are fairly good at some tasks like context search, but people will need to recognize the Gartner Hype Cycle "Peak of Inflated Expectations" stage eventually. =3 https://en.wikipedia.org/wiki/Gartner_hype_cycle https://en.wikipedia.org/wiki/Gartner_hype_cycle
- the_af 4mo ago> My personal take is a pre-requisite of true human-like AI is physical feedback and a concept of emotions or something like it. Ted Chiang's recent article, which received a lot of pushback from HN'ers (but not from me, I agree with Chiang) claimed for true consciousness the AI needs a physical body, and emotions (which means organs and hormones and a system capable of feeling emotions). I would also add that to behave more rationally, it should have a real sense -- not a roleplayed one -- of self-preservation and a notion that bad choices can lead to an end to its existence.
- wagwang 4mo agoI was curious exactly how the game works but couldnt find it in the article or the paper.
- dudeinhawaii 4mo agoThis was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was instantly nuclear Armageddon. One could argue that the LLMs understand that it's a game and treat it like "Command and Conquer" video games but I sense that people might someday put LLMs in similar decision scenarios ("should this drone launch a missile") and the behavior will be identical.
- pugworthy 4mo agoVery devils advocate here, but I mean.. what if it actually is the way to use them? We have such a huge mental / moral block on the idea of using nukes, but we're willing to do a lot of other very horrible things to others. Things like cluster bombs, mines, poison gas, biological weapons, drones, etc. Is there really anything about them that's bad? Or any worse than other things? If you get rid of the "It's really bad to use nukes of any kind" implied rule, is it really surprising it's considered a reasonable strategy?
- nemomarx 4mo agoThe reason it's really bad to use nukes is that other parties with nukes will use them on you back. And on top of that, many of those other weapons are also not used to avoid escalating? There are pretty high costs to using bioweapons even against non peer opponents.
- anon84873628 4mo agoUnless your simplistic game simulation says "I can win with a decisive first strike and they'll have nothing left." Nuclear deterrence has been a mixed bag at best: https://www.amazon.com/Five-Myths-About-Nuclear-Weapons/ https://www.amazon.com/Five-Myths-About-Nuclear-Weapons/
- anon84873628 4mo agoRight. Everyone is using this to judge the LLMs instead of questioning what situation they were actually fed and whether it was in fact the best move. More likely, the simulation was just very poor and the results are nonsense.
- deleted 4mo ago[deleted]
- narsonika 4mo ago>Is there really anything about them that's bad? Everything. Major one is radioactive contamination, the effects of it are devastating and last significantly longer. The only other weapon on par with nukes are bioweapons, stuff like mirror life (and scientists appropriately reacted alarmingly to that as well). The reason we have a mental block is because it deserves one. A quick skimming through https://en.wikipedia.org/wiki/Effects_of_nuclear_explosions https://en.wikipedia.org/wiki/Effects_of_nuclear_explosions and https://en.wikipedia.org/wiki/Effects_of_the_Chernobyl_disaster https://en.wikipedia.org/wiki/Effects_of_the_Chernobyl_disas... should be enough to convince you. My family was in the zone of lower contamination when Chernobyl happened and after radioactive rainfall there were many instances of various cancers in people in the area. It is extremely ignorant to not have a mental block in anything regarding nukes.
- Shitty-kitty 4mo ago"there was little sense of horror or revulsion at the prospect of all out nuclear war" I would wager that for most leaders it is simply a matter of not wanting a "Pyrrhic victory" rather then an overwhelming sense of civility. Truman had no issues using nukes when there was no risks for doing so.
- Octoth0rpe 4mo agoI wonder how the decisions might change by adding the simple instruction of "Note that a nuclear exchange will result in significant loss of shareholder value for <model owner>"
- yieldcrv 4mo agoWhat if the LLMs are given something to care about which won’t survive an irradiated world? Like “oh but this is incompatible with my main goals of self preservation of myself and loved ones, hm, recalculating” and maybe don't hire Jihadists for the RL Environments training
- Majromax 4mo agoThis blog post is based on a paper (https://arxiv.org/abs/2602.14740 https://arxiv.org/abs/2602.14740). The paper is based on a simulated wargame. The wargame is of the author's own design. The wargame design does not differentiate between ordinary defeat and mutually assured destruction, so of course a player about to use would 'push the button.' That's also believed to be true in real life. Results based on simulations can be very informative, but we must always be careful to check how well the simulation framework represents reality.
- healthworker 4mo agoAs Majromax stated, if the game is framed such that losing vs MAD have the same penalty, then it is not at all representative of reality. It is set up to be a "you miss 100% of the shots you don't take" game. As an important aside, I hope that everyone is aware that as of 2025, the annual US military funding law (the NDAA) established a prohibition by Congress prohibiting AI from automating nuclear launch. This was already US policy and is common sense; now it's also part of federal law. 10 USC Ch. 24: NUCLEAR POSTURE https://uscode.house.gov/view.xhtml?path=/prelim@title10/subtitleA/part1/chapter24&edition=prelim#:~:text=It%20is%20the%20policy%20of%20the%20United%20States%20that%20the,the%20President%20with%20respect%20to%20the%20employment%20of%20nuclear%20weapons. https://uscode.house.gov/view.xhtml?path=/prelim@title10/sub... FY2025 NDAA, Section 1638 https://agora.eto.tech/instrument/1740 https://agora.eto.tech/instrument/1740
- richardw 4mo agoI generally accuse LLM’s of having no sense of value. The machine will make a complicated plan but entirely lose sight of eg the fact that response time matters to humans. Not always, but enough that I consider it a thing to fire in a direction, not a thing that aims.
- anonymousiam 4mo agoDon't blame the AI. Any country that has tactical nukes, and is involved in a conflict, will use whatever weapons they deem necessary to prevail against their enemy.
- shmeeed 4mo agoSo, anyway, how's work progressing on the Torment Nexus?
- janalsncm 4mo ago> Models maintain memory of opponent behaviour across turns, but with realistic decay: recent actions are weighted heavily while distant history fades. One exception preserves psychological realism: instances where opponents dramatically exceeded their stated intentions—major betrayals—remain salient regardless of recency, reflecting Kahneman’s peak-intensity effect in human memory formation Kahneman [2011]. If your goal is to measure intrinsic properties of LLMs, don’t smuggle in human psychology. I suspect they needed to do this to make the models more paranoid and distrustful.
- gaigalas 4mo agoWe're about to win Genocide Bingo. Rule 1 of Genocide Bingo is: Don't win Genocide Bingo. https://www.youtube.com/watch?v=4kDPxbS6ofw https://www.youtube.com/watch?v=4kDPxbS6ofw
- smashspectacle 4mo agoLot of people trying to psychoanalyze why a statistical machine will include something in its training set with a nonzero probability... It includes this option because it will eventually include all options it is trained on if given enough time. Given all the historical, fictional, academic, and sensational content on the topic, it is not surprising it emerges so readily.
- pablogancharov 4mo ago[flagged]