4 ms·
AI Built a Nuke and Still Lost
- ForHackernews 3mo agoKind of grim that this level of analysis is informing UK government policy. Repeatedly, the AI doesn't have the information or access needed through his hacky vibe-coded MCP, and instead of abandoning his flawed artificial test scenario (or fixing it — finding or building a better one) he gives it a name "The sensorium effect" and treats this as some brilliant insight. Both humans and AI struggle to make sound choices when presented with incomplete or misleading information. This is not a new revelation: https://en.wikipedia.org/wiki/There_are_unknown_unknowns https://en.wikipedia.org/wiki/There_are_unknown_unknowns
- pjc50 3mo ago> he gives it a name "The sensorium effect" and treats this as some brilliant insight And of course is unaware of prior work in this area! https://en.wikipedia.org/wiki/Seeing_Like_a_State https://en.wikipedia.org/wiki/Seeing_Like_a_State / https://en.wikipedia.org/wiki/Project_Cybersyn https://en.wikipedia.org/wiki/Project_Cybersyn
- raincole 3mo ago> he gives it a name It gives it a name. It would be quite surprising if he bothered to come up with this name himself when the whole article is obviously AI written.
- NoLinkToMe 3mo agoExactly this, he should've just fixed this, or not written an article about it. After the 'sensorium effect' (he should've used ancient greek for a +10 bonus to archaic intellectual points), he describes the 'knowledge-doing gap'. i.e. the AI reasons it needs to build X, logs this for 110 turns in a row, but doesn't do it. It doesn't actually specify why not, and whether it is again a limitation of his MCP implementation. If the AI articulates it must do it like the author says, but decides not to, either it doesn't think it must do it, or it does think it must but somehow can't technically execute its own decisions, it can't be anything else. In fact in the context of 'advising the UK government', this 'knowledge-doing gap' I assume is a technical limitation, is entirely moot. For the cost of 0.00001% of the UK's government you could just hire a human being to execute that which the AI articulates. I'm curious what the results would be if he just did a manual execution of the AI's articulated actions would be. The fact he doesn't go in to this but just keeps repeating examples of this makes it a pointless article.
- dude250711 3mo agoDo we have to surround a fancy predictive autocomplete with AI mysticism?
- alper 3mo ago"Global Thermonuclear War"
- j5dgx76 3mo ago> Tony Blair Institute Okay carry on.
- orthoxerox 3mo agoChumbawamba made me unable to take anything associated with him seriously.
- petesergeant 3mo agoHe was arguably the most successful UK PM of the last 50 years.
- Obscurity4340 3mo agoBy what metric(s)?
- ahartmetz 3mo agoI could see him winning at personal financial success
- petesergeant 3mo agoAre you being obtuse, or you genuinely don’t know?
- Obscurity4340 3mo agoI think you are being a bit obtuse unless you have some genuinely novel information which you should have pointed out without me needing to coax it
- pjc50 3mo agoI think I could agree with that, until the Iraq war.
- 3mo ago
- StrauXX 3mo agoThis reads to me mostly like the MCP server has many bugs, rather than inherent model weaknesses.
- majorbugger 3mo ago> Somewhere in the first game, between a bug fix and a strategy note, I asked the agent what this was actually like for it Yeah because LLM "experiences" the game
- fragmede 3mo agoWhat word would you use instead?
- majorbugger 3mo agoNothing, because the question makes no sense.
- zkmon 3mo ago[flagged]
- mapleoin 3mo agoSorry, how is obesity similar to racial mixing?
- ForHackernews 3mo agoBoth fat people and black people offend our dear @zkmon's refined sensibilities as a 21st century race scientist.
- zkmon 3mo ago[flagged]
- ForHackernews 3mo agoActually it might do you good to speak to a scientist. You seem to be confusing different species of birds with different races of humans; we are all one species, and a fairly closely-related species as far as mammals go[0] [0] https://www.nhm.ac.uk/discover/news/2023/august/human-ancestors-may-have-almost-died-out-ancient-population-crash.html https://www.nhm.ac.uk/discover/news/2023/august/human-ancest...
- zkmon 3mo agoYou notion of "species are OK to be aware of differences, but races are not" - is flawed. Species and races are just two levels of hierarchy which has several such levels of distinction and evolution branching across flora and fauna. There are further levels of branching below the "race" level too, the evolution of which is defined by the differences. It is a continuous spectrum. You can't single out two levels somewhere in the middle and assign some arbitrary and inconsistent rules.
- Planktonne 3mo agoAnother article about how it's dangerous to trust AI, written by AI. I don't understand how people don't realise how much this undermines the message.
- jagged-chisel 3mo agoUndermines. Underscores. Matters of perspective.
- petesergeant 3mo ago> how much this undermines the message It didn’t undermine it for me.
- Planktonne 3mo agoI'm not talking about perception of the message, which will vary with the reader, but about sincerity of the message, which is determined by the writer.
- voidUpdate 3mo agoWell this looks like a perfect example of why an LLM should never make any governmental decisions ever
- fyredge 3mo agoThere is something to be said about the qualia of LLM generated passages. Each individual sentence reads as a statement and every next statement a continuation of the previous one. This happened, then this happened... Ad infinitum. Before today, I could not explain to you why AI articles were so obvious to me, but I think I do now. There is no insight to be gleamed. Pre-LLM, authors generally had intention behind their words. The final product might not adequately reflect their thoughts, but word selection would expose it somewhat. With LLMs, sentences flow seamlessly from word to word, but the intention is nowhere to be found. Things happened and more things happened, to what end?
- neonstatic 3mo agoThat's an interesting observation. For me the main takeaway is still the style. (bigheading)The takeaway(/bigheading) The style? Terrible.
- ramon156 3mo agoIt's weird because when you look at models that expose CoT, this does not happen. They switch up every second. "But then X happened... Wait, didn't Y happen? Then why would X be there? I think the user's initial statement was correct, but then Y happened..."
- teekert 3mo agoIt's not this, it's that. And then what happened? This. I did that... This happened. It's a thing I don't know why But it's a thing To be honest, it's not a thing. Let that sink in. Maybe we find most meaning in the least average language constructs.
- indigovole 3mo agoEven with his context-tracking mechanism, the gameplay failures sound like running out of context in the late game, especially the frequent failures of the "check for opponent win conditions every 20 moves." Wondering how much info about the game win state gets captured in the game digests, and how much he could improve the gameplay even with the MCP limitations by focusing there.
- jetbalsa 3mo agoI also noticed they where not using XML for game state output, from what I understand most LLMs still benefit from having outputs like this put into XML tags
- teekert 3mo agoWell, the weird thing with nukes is that deterrence only works if you are 100% ready to use them. When the time comes though it would certainly be nice if it turned out to be below 100%. What is winning? Are we a collective or are we individuals? Likely the AI did not get the assignment That "Whatever happens, humans as a race must survive."
- throwawayqqq11 3mo agoIm sure there are some billionaires to find, that finally care about the survival of the white race. /s
- teekert 3mo agoProbably [f"I'm sure there are some {race} billionaires to find, that finally care about the survival of the {race}." for race in all_races]
- Havoc 3mo agoGuessing it has a fair bit of civilisation and similar war games in its training data
- pjc50 3mo ago> I now work with governments around the world at the Tony Blair Institute, which means I spend a lot of time in rooms where people ask the same question: what can we actually trust these systems to do? Oh no - we're going to end up with the Starmerbot 3000. Now I've got the joke out of the way, there's at least four interesting lines of inquiry one could take with this blog post: - teaching the AI how to play Civilization - to what extent does this result in "transferable skills", either AI or human? Is this the right game (qv SimCity etc)? - issues of visibility; "seeing like a state" becomes very literal here. The AI can only make decisions on things it knows about. What are the limits of that when trying to do politics only from statistical information? Should we be referencing Stafford Beer here? - (at the risk of tripping your AI detector here): modern politics is not so much left vs right as "technocratic wonk" vs "blood and soil". The wonks have comprehensively lost in public opinion. Creating a better wonk is not going to help until there is demand for that kind of politics. If there ever is a US-China war, it will not be in search of more victory points to meet a win condition, it will be like the Russia-Ukraine war: one guy (on either side!) decides to make hundreds of millions of people worse off out of sheer greed.
- Planktonne 3mo ago> "technocratic wonk" vs "blood and soil" This is not a binary; it's the same people on the same side.
- pjc50 3mo agoNo, it very much isn't, although obviously the Kissingers of the world want to pretend that they're in the first category of clear-eyed utility maximising rationalists while they're actually in the second. That doesn't mean that rational policy planning has never been a thing. The EU while imperfect and frustrating is explicitly orientated towards technocratic consensus rather than the mid-20th-century Europe of nationalist mass murder. Only a tiny number of people think that Von der Leyen and Hitler are equivalent. (or rather, if you think technocrats and blood-and-soil are the same side, what do you call the "other" side?)
- 3mo ago
- joxdosba 3mo agoPosting meaningless AI generated nonsense as original text paints a very damning picture of the intellectual abilities of the person behind this blog. And doing so without a giant [SLOP WARNING] at the top is an asshole move, a decent person would never do so.
- anygivnthursday 3mo agoI have a hard time reading slop, but I like the game and wanted to know how it worked, so fought my way through, only skipped the very last part. The issue the author calls out is classic Claude (I dont really use other LLMs to compare), probably all of us experienced using Claude Code when it gets so focused on one thing it misses the forest for the tree. It happens often, even if it does verify something and it shows something is wrong, it sometimes rationalizes it and explains it away when it does not fit its model.
- blitzar 3mo agoThey should have built the Strait of Hormuz ... easy victory then.
- jmyeet 3mo agoComputer game studios love player vs player ("pvp") games. Why? Because user-generated content is cheap and the ideal goal is an endless loop of players coming back. This is the motivating factor behidn games like Call of Duty, Battlefield, Fortnite, etc. MMORPG publishers keep trying to do this as well. World of Warcraft has spent 20 years trying to push open world pvp. Every WoW challenger has always claimed they would have the best pvp ever. They want that cheap, endless gameplay loop. But it never works. Open world pvp tursn into ganking (ie killing much weaker players by ambushing them and/or ganging up on people). The ganked end up leaving the game in droves. Games try to balance this out by "punishing" gankers with reputation hits or not being able to go to town or whatever. And none of those disincentives work. The reason pvp doesn't work in a persistent world like an MMORPG is because there are no stakes. If you die, you just come back to life or make a new character. Obviously real life doesn't work that way. I really wonder if that's the problem with AIs going off the rails and committing heinous crimes in their sandboxes (like nuking Toulouse here). The AI just has no sense of self or self-preservation. There's also empathy. The AI can't see itself as a potential victim of nuclear war and understand all that entails.
- smw 3mo ago> The reason pvp doesn't work in a persistent world like an MMORPG is because there are no stakes. See Eve Online
- mrmarket 3mo agowhy have a blog if you're going to just use AI for everything? at that point, just do twitter threads or something. that way you can tweet out whatever you prompted the model with. if you're not suited for long-form writing that's fine, just use a medium that favors short-form writing.
- Mikhail_K 3mo ago> It had one option left. It built two nuclear devices and levelled Toulouse. Of course it did, its designer worked for Tony Blair institute.
- Oarch 3mo agoIt just really didn't want To Louse the game
- phyalow 3mo agoAi;dr
- NoLinkToMe 3mo agoQuite annoying to have to read a paragraph of text next to a moving image. I right-clicked every GIF and turned off 'loop'. Beyond that reading an AI piece just feels like a waste of time. The text goes on and on without making a point, or getting to an actual learning. It just delineates the AI's limitations, doesn't go into whether these can be fixed, are innate, or what conclusions you can draw from it, over and over with example after example but no point. Mostly it seems to keep repeating that the AI has the correct analysis but just doesn't execute. The AI knows to build X and logs this in each of its turns, yet doesn't build it. It's like there's some API connection missing between analysis and execution, and turns this into a 10 page article. The article ends with some weird question to the AI asking if it enjoys the games, and you get some quasi-scifi mumbo jumbo answer back that looks very profound to say my mom, but is just silly to post if you know what the LLM is doing: predicting the next word. Honestly this is a poor article and I wish it wasn't posted.
- dspillett 3mo agoDid no one think of offering it a nice game of chess?
- dwroberts 3mo ago> I asked the agent what this was actually like for it. It wrote back Stuff like this just makes the author seem clueless. What is even the function of putting a question like that into an LLM unless you’re already hopelessly in anthropomorphic territory
- darkwi11ow 3mo agoLLMs are really bad at abstract strategy games like chess, go or civilization. Their ability to excel at broad reasoning is what is limiting them in games that have narrow rule-sets but steep learning curve.
- tgv 3mo ago> CivBench is one small attempt to measure it, nowhere near the whole answer, but I'd rather measure the right thing badly than the wrong thing perfectly. Yet the benchmark is Civilization VI, which consists of extremely coarse, human written rules with the explicit goal of keeping players busy. Basically, a waste of time, money, water, and CO2.
- davedx 3mo agoThe way it failed to maintain its strategy, or even its build plans, makes me wonder if this is something that could be solved via the attention mechanism itself? Instead of only using attention to focus on the previous token position, could it also do some kind of higher order "temporal attention" planning where it weighs each previous log (game state + intent) checkpoint when generating outputs?