4 ms·
Brood War Bench
- aswegs8 13d agoCheck out pluto, RL trained sc bw bot. Somehow it is now going rogue on the Korean ladder and flattening pros. There are some yt games by it, search for ^333^
- benswerd 14d ago+ Playable Agent driven Starcraft
- conorcleary 14d agoDibs on the fly brain
- indrora 14d ago"THIS FLY CAN BEAT YOU AT STAR CRAFT: HAS SCIENCE GONE TOO FAR????"
- bee_rider 14d agoA ton of conversations about the game must be in the training set. I wonder, is there any way just from watching how they play, of telling if they tend to pick strategies that people complain or meme about online?
- stackghost 14d agoI wonder if there is a library to decipher brood war replay files. Perhaps an agent could learn by watching.
- frutiger 14d agoI haven’t checked for SC:BW but Blizzard has official parsers/replayers for SC2 on GitHub.
- benswerd 14d agoNot hard to build. I was shocked at how fast/easy this was to pull together.
- duskwuff 14d agoYes, e.g. https://github.com/gulshngill/bwrepanalysis https://github.com/gulshngill/bwrepanalysis. But I doubt an agent would learn much from them without a lot of additional processing - the replay files are little more than a stream of the orders given during the match (e.g. "at tick 17, player 1 ordered unit 234 to attack-move to 56,78"). They're difficult to make sense of without a lot of additional context, like the map layout, the location and status of other units, what parts of all that are actually visible to each player, etc.
- tekla 14d agoAt high levels of play Zerg is generally considered significantly stronger than the other races. So I generally will assume AI will tend to pick Zerg
- xmcp123 14d agoUnless you Protoss and mind control the Zerg, and then are Zerg+Protoss.
- duskwuff 14d agoThis is not a viable strategy in competitive play. It's a huge resource/time investment with a dubious payoff. If you could win as Protoss by building Dark Archons, raiding your Zerg opponent's base to capture a worker, building a bunch of Zerg production and tech buildings, and attacking with a combined army... you most likely outclass your opponent, and could have won much faster using a more conventional strategy.
- xmcp123 13d agoTo be fair I’m recalling 2v2 NR20 games from an era where I played talking to my childhood friend on a landline. So there was enough time. The big risk was that the map would max before you could fully develop the Zerg side.
- duskwuff 13d agoWhat you're describing is a casual game, not competitive play. And, of course, all bets are off if you're playing with additional rules like "no rush" which remove the element of tempo from the game.
- rrr_oh_man 14d agoI only played the StarCraft demo two decades ago. Why is that so? Rushing?
- 14d ago
- malfist 14d agoThis is a great idea for a benchmark. Something all the benchmarks seem to be missing is strategy, tactical solutions in most of the benchmarks are all thats required but here requires actual long term thinking and tactical thinking, balancing and orchestration.
- shard972 14d agoThat’s why I tried making https://wrathbench.shard.page https://wrathbench.shard.page
- tyre 13d agoWow this is super cool. I wonder when LLMs get cheap enough that NPCs will be able to have full conversations with players. Not sure they’re the best option for raiding, but as a high-level orchestrator for choosing content, that sounds pretty great.
- stymaar 13d agoJust curious: do you know the expected XP gain over the first 90 minutes for a human player (either a medium-level one or a speedrunner)?
- deleted 14d ago[deleted]
- Game_Ender 14d agoAny details about the harness the agents were given? I am curious what representation of the screen and world state was provided to the agents and what tools they had available.
- benswerd 14d agoOh sorry I should be more clear on that. Will add to report. For agent harness I did Claude Code, Codex, Grok Build. This was primarily a cost driven decision — I have a lot of free tokens and I didn't want to pay API prices for this. For game harness I used minimal BW-API issue command and get observation apis as tools. I felt this was the most fair way to do it on my small scale. In the future I would like to integrate code mode and multiple games/I think if it was a best of 5 where each agent could learn from its past games and build its own automations over time that would be much more interesting.
- usef- 14d agoGiven that a lot of their failures are from fairly basic mistakes related to the unique setup (eg, thinking rather than defending immediately) I'd love to know how much they improve with basic tips. Or possibly even whether they can learn from a game themselves. "Analyse your game for your failures" -> Then give a fresh agent of the same model that "learnings" doc for the next match. Do the rankings change over time, if models can write instructions for future selves?
- benswerd 14d agoTry it, you can run your own games on bw.swerdlow.dev
- pelagicAustral 14d agoUnrelated to the benchmark... I love StarCraft. I started playing it right from the beginning, most of my friends right now are from that era. I literally met people that have spread to almost every continent when I was in my early teens. We played at internet cafes and did not have access to the internet, that was priced differently... I miss those days so much. Everybody was from a different background back then, and nobody was anything other than a guy that plays StaCraft at the cybercafe... And now, we are in our 40's and I know Math teachers, history teachers, oil rig operators, software programmers, professional gamers, lawyers and more... hahah So crazy to think about it... and I know them, we talk, what a world.
- sidewndr46 14d agoLater we even listened to Eminem and played violent video games. Most of have never even been charged with a crime, much less abused someone.
- benswerd 14d agoI agree. My first time playing StarCraft was at summer camp around a decade after it came out. All the smartest people played it so I wanted to too. Great decision, I have been continually impressed with the people who StarCraft introduced me to.
- nemo1618 14d agoEven at Burning Man, in the middle of the desert, there is a camp that hosts a StarCraft tournament every year (on the dustiest setups you've ever seen!) :)
- benswerd 14d agoSo dope
- PorciiVorbesc 14d agoSo ... Zerging Man?
- jpgvm 14d ago
- GodelNumbering 14d agoA friend of mine created GoBench[1][2] that evaluates LLMs on 9×9 Go using KataGo opponents as Elo anchors, you see real capability differences there, like Astra Max substantially leading all other models. I think strategy is a generally interesting area to evaluate LLMs on [1] https://rolandgao.com/blog/gobench/ https://rolandgao.com/blog/gobench/ [2] https://rolandgao.com/gobench.pdf https://rolandgao.com/gobench.pdf
- nullc 14d agoA better benchmark might be asking the LLMs to write GO ai and then comparing that-- the issue is that there will be a HUGE difference in performance that depends purely on this game being in the LLM's training... but training a general LLM to directly play these games would be a waste of capacity and shouldn't be encouraged for benchmaxxing sake. Programming an engine OTOH is a skill that is more general and they should all have. Might be useful to have the target of the engine be some specific virtual machine that gets a strict cycle budget-- e.g. execution runs so many cycles, and result is read out of a specific memory address at the end (or when it terminates early).
- ForHackernews 14d agoExcuse me? Is this thing supposed to be a borderline SGI or not? We already know LLMs are good at spitting out code.
- nullc 14d agoPlaying these games autoregressively isn't even the right way to use the LLM for this task (unless it was trained to do so...). It's somewhat like having a creative writing bechmark but requiring that all the input/output be base64 encoded. It can do it-- but no guarantees on the results! And it's also just bencmaxxing bait: you can get a huge improvement on the task by RLing on it, but make no improvement on anything else. Doing so would just waste model capacity. If you could tell that every LLM was equally not being exposed to the task then you could justify it as a test of abstract reasoning, but you can't. So it ends up on how much go transcripts ended up in the training, which is ... not a very interesting metric.
- winwang 14d agoWould be interesting if you could team a fast and slow agent together -- slow model can either act directly or maybe just communicate to the fast model.
- benswerd 14d agoI might open this up to a tournament if enough people want. Any interest?
- winwang 14d agoHah, yeah that's definitely interesting, though maybe a general platform for this stuff would be even more interesting. Although, maybe benchmarking an agent on "how well can you command a swarm to annihilate the Terrans" is how it all starts going downhill...
- dschuessler 14d agoSomewhat related: In 2018, Google DeepMind had already created AIs that were capable of beating professional gamers in StarCraft 2 (the sequel to Brood War): https://www.youtube.com/watch?v=cUTMhmVh1qs https://www.youtube.com/watch?v=cUTMhmVh1qs
- benswerd 14d agoI predict LLMs will reach superhuman level and beat even that model in the next 12 months
- orbital-decay 14d agoStarcraft is APM-dependent. Unless the latency will improve greatly in frontier reasoning LLMs (which is unlikely), it will remain a bit like knitting with an excavator.
- gadtfly 14d agoDid it play by looking at screenshots and sending clicks, or was there other mediation/symbolization? It sounds like it might have been actually played in real time, which would be very important to distinguish. I have recently seen other harnesses letting agents play real-time games in what seems like discrete time slices, turning eg Portal into something turn-based https://www.youtube.com/watch?v=ruuGXFAmiOE https://www.youtube.com/watch?v=ruuGXFAmiOE
- loeg 14d agoThere's a bot data stream already; it's probably hooked up to that rather than screencap. Yes, I believe these were playing in real time.
- callmekit 14d agoIt plays from a special API. Some things that are impossible with normal UI are possible with API, like selecting a unit under other units. Invisible units are also reported via the API.
- faeyanpiraat 14d agoThere is currently a bot beating everyone on the ladder. Just watched it today on Artosiscasts yt channel.
- ericpruitt 14d agoLink to the video for the curious: https://www.youtube.com/watch?v=1vsTqNwHquE https://www.youtube.com/watch?v=1vsTqNwHquE . I would note the bot is not APM limited, and comments suggest (https://www.youtube.com/watch?v=1vsTqNwHquE&lc=UgxBrY4CkuZx7oy5dex4AaABAg https://www.youtube.com/watch?v=1vsTqNwHquE&lc=UgxBrY4CkuZx7... , https://www.youtube.com/watch?v=1vsTqNwHquE&lc=UgyZ0XVFDqzJTgmiBN14AaABAg https://www.youtube.com/watch?v=1vsTqNwHquE&lc=UgyZ0XVFDqzJT...) it's likely using map hacks, both of which make the bot significantly less impressive.
- vitaflo 14d agoIf you read the TL thread on this, the bot is technically running at 240 APM, but its commands get split off into separate commands for each unit, so the APM looks higher. I've watched several replays on this. It using map hacks is lame but it also doesn't always take advantage of them. It sometimes does respect its own fog of war. It mostly wins with incredible micro (kinda has to, it's macro kinda sucks). Most of its few losses come from drops (it never makes turrets and doesn't know how to handle them), bad macro (blocking its own ramp) or just incredibly unconventional play from its opponent (which is certainly not in its training set).
- chrishare 14d agoSpecifically, it can micro across multiple screens of engagement in a way that we can't as easily.
- stephbook 14d agoStarCraft is just not a good benchmark exactly for these reasons: Endless complaining about the bots not being limited in the way humans are. "Oh it's able to issue commands too fast." "Oh no, I give it full map access and it uses that." Ego shooters are, naturally, also not good benchmarks. OpenAI won DotA2 in 2019, a way better game.
- tweakimp 14d agoIf you want to see human written bots in action or compete in the bot ladder yourself, try https://aiarena.net/ https://aiarena.net/
- minimal_action 14d agoI think we're on the early days of games you connect with your agent to. Human + AI units one versus the other. Like knights with their horses. Not sure which is the horse..
- WillMorr 14d agoI've been running these with friends recently, it's very fun. I ran irl bot tournaments for board games a couple times but the agent era opens up a huge amount of possibilities. Must recently I built out a MMORPG puzzle box thing, I wrote a general game architecture doc but left the specific puzzle design up to Fable. Nobody is actively playing rn but I left it up at bot.willmorrison.net.
- AntiRush 14d agoBack in 2010, during the early days of bwapi, there was a Brood War AI tournament held by the Expressive Intelligence Studio at UC Santa Cruz. It's interesting to see how different the approaches were back then, vs this or Deepmind's SC2 work. https://web.archive.org/web/20091124210529/http://eis.ucsc.edu/StarCraftAICompetition https://web.archive.org/web/20091124210529/http://eis.ucsc.e... There's a great contemporary Ars Technica piece by a competitor: https://arstechnica.com/gaming/2011/01/skynet-meets-the-swarm-how-the-berkeley-overmind-won-the-2010-starcraft-ai-competition/ https://arstechnica.com/gaming/2011/01/skynet-meets-the-swar... As an undergrad I did a project using genetic programming. It was not very successful, but it was a lot of fun. https://tomisin.space/archive/starcraft-genetic-programming/ https://tomisin.space/archive/starcraft-genetic-programming/
- dcl 14d agoIn the early days of SC2, I remember people using genetric programming to optimize build orders. I remember a slightly unorthodox Zerg Roach Rush which was _really_ fast.
- tinco 13d agoAll the hobbyists were using proxybot because it allowed you to use more fun languages than C/C++ but proxybot lacked the features to effectively play Zerg. I really wanted to play Zerg so I built a really sweet API+DSL in Ruby around proxybot and then used that to give the other newbies a hard time with a zergling rush. Unfortunately the proxybot limitations precluded me from expanding its capabilities so I tried to switch to having ruby embedded in C++ and I basically got mired there and was eventually distracted by real world concerns like actually finishing my degree. I think I played against Krasi0's bot a couple times in the early days. Hopefully they stuck with AI and have suddenly become crazy rich after 2016. It certainly wasn't a given that AI was going to lead to a fruitful career back then, let alone to unimaginable riches.
- mslate 13d agoI also entered the competition as an undergrad—I contacted Ben Weber 10 years later (2020) and interviewed him on my podcast: https://theaccidentalengineer.com/adversarial-machine-learning-ben-weber-zynga/ https://theaccidentalengineer.com/adversarial-machine-learni...
- c7b 14d agoAt last something that feels properly orthogonal to pelicans on bicycles.
- chaostheory 14d agoGemini wasn’t included, but I’m guessing its performance would have been similar to Grok’s performance despite having roots in DeepMind.
- mococa 14d ago+1 because StarCraft
- moomin 14d agoWondering what it would look like if you allowed them to write scripts. You could throttle the number of clicks to make it interesting.
- therealdrag0 14d agoWould Jev be good for this?
- aetherspawn 14d agoIt can play doom in realtime so maybe it can play this.
- mcteamster 14d agoI love this. Funnily enough StarCraft has influenced how I approach AI at a meta level Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks Terran: versatile team comps of dedicated agent roles you can delegate well-defined tasks to Zerg: massive swarms of specialist custom agents inside your apps that you evolve and optimise for speed and cost Knowing every faction has its strengths and weaknesses helps me decide which tools to use for the job.
- alchemism 14d agoOn a tangential note, try run Vibe Island (or equivalent) with Protoss WAVs for the agent sound effects.
- mcteamster 14d agoCommitted to the bit; context aliased “change plans” to “pylons” and I have a command to /construct them
- alchemism 13d agoEn taro Adun, Executor.
- Barrin92 14d agoGame performance is one of those topics that makes it so abundantly clear how limited these systems still are. StarCraft is predominantly a mechanical game so the horizon of what you need to do is quite short and tactile, and even then without advantages no system has come close to beating a human. I saw someone recently try to get an agentic system to play Final Fantasy and it did about as well as a Roomba.
- ethanpailes 14d agoAlphaStar was absolutely dominant. I don’t know if it played exhibition matches after it got its view restricted, but surely it still performed at a very high level. The APM restriction they put on is actually a handicap in favor of the human since humans are allowed to have 2k APM, just incapable of doing so. If you just mean general LLMs can’t beat humans yet that’s one thing, but it’s not the case that no system can do so.
- Barrin92 13d agoAlphastar played with cheats such as full map vision, which obviously in starcraft makes a gigantic difference
- suby 14d agoI don't know where else to write this, but I want to throw the idea out there. I have long wanted to take old broodwar televised matches, many of which are terrible quality 240p, and use machine learning to convert them to into perfect Broodwar Remastered frames. This seems tractable to me because you should be able to map the terrain sets to their remastered equivalents, and the game is just a series of sprites rendered at specific frames. Even if the source quality is terrible, I imagine this is able to be extracted at high quality since you can, eg, set up an automated pipeline which generates training data. Maps from original graphics to remastered, and then again for 240p -> tilemap positions for frame camera center + sprite positions / animation index.
- silentkat 14d agoLink to the matches?
- sharkjacobs 14d agohttps://www.youtube.com/@CholeraSC/videos https://www.youtube.com/@CholeraSC/videos
- efxhoy 14d agoI would love all the NukeTheStars matches in 4k, his commentary was fantastic
- KeplerBoy 13d agoMaybe one could try to reconstruct very closely matched replay files?
- 3eb7988a1663 13d agoDo replay files require consistency or can you spawn new units from the ether as required? I think you could get a system that would get the broad strokes unit placement correct, but having a shot-for-shot perfect replication seems impossible.
- DJMolehill 14d agoworking on a similar project for street fighter currently https://www.youtube.com/watch?v=dJYTV3ZiT6o https://www.youtube.com/watch?v=dJYTV3ZiT6o
- windowshopping 14d agoI would love to create one of these benchmarks for age of empires 2, but I have no idea how to make the AIs play it. Maybe I can get Claude to do it anyway.
- xyzsparetimexyz 14d agoGood example of how these things sometimes spend way too much time thinking to be useful
- American87 14d agoI don't normally anthropomorphize the AIs but this is super cute, lol.
- Svenstaro 14d agoI'm trying to play but it says "another match is already active". Can only a single player play on your system at the same time? EDIT: Nevermind, seems to work now. Watching a local qwen3.8-flash-next play this.
- snikeris 13d agoYeah, what's the deal with this error message?
- deleted 14d ago[deleted]
- bigcat12345678 14d agoOpenai 5 dots 2, and alpha Star, AI research used to be very fun
- kevinrineer 14d agoAlphaGo, Watson, Stockfish, Eliza. I can name so many of the interesting stuff and each new LLM just comes and goes.
- rubiquity 14d agoNo model will ever figure out how to make a ling tight wall.
- karim79 14d agoThis thing kept me sane through university. I had a shitty computer which could barely run this and no Internet so I just played against the computer, which was both frustrating and educational. I will always love this and now I'm going to play it again. Remastered and on a fancy modern machine.
- karim79 14d agoYour minds will be blown when you realize just how much StarCraft is ingrained into South Korean culture. They literally had (or have) dedicated TV channels just for StarCraft. Brood War was the first video game to be broadcast on TV in Korea. I'm pretty sure it's still going.
- vanderZwan 14d agoEh… I'm fairly sure that most of the people here are old enough to have played Star Craft and Brood War when it first came out, and are fully aware of e-sports becoming huge in South Korea. And I have a hard time imagining younger generations that grew up with e-sports being an established thing instead of a novelty to not be aware of how South Korea looks at Brood War differently.
- iririririr 14d ago*wonders if running startcraft in wasmjs uses less CPU than anubis
- stymaar 14d agoInteresting that it benches down to Haiku but doesn't bench any Chinese models (which are at least between Sonnet and Opus, when they aren't beyond Opus).
- mkotlikov 14d agoWhy was Luna Low so good?!? Better than Terra XHigh. Can't even say it's all APM because Sol Medium beat Sol Low (both beat Sol XHigh).
- ashdnazg 13d agoMaybe just speed? TFA said that many agents failed because they were thinking too much and doing too little.
- sqrt_1 14d agoI thought that it was videos of the replays on the site. Was very impressed that it was a replay playing that you can scroll around in and select units.
- Narishma 13d agoIs that why you can't pause them? I just ended up using uB to zap them since they were taking too much resources.
- agentdev001 13d agoYou must construct additional pylons
- Karliss 13d agoYour APM is too low. If you click a bunch sometimes the pause button works.
- Havoc 13d agoI like this as a concept - taking a very real world task and checking whether it works. Surprised the outcomes are so poor though. I recall years ago AI was capable of beating pro level DOTA teams. I guess in one case it was specifically trained on the interface & game while here it was not?
- mrkeen 13d agoGame programmers have long been the producers of the most impressive applied computer science. The film Shrek 3 (2007) took 20 million CPU hours of render time. Games push 60 frames a second. For a visual comparison, check Call of Duty world at war (2008). Also compare to browsers, which can sometimes scroll smoothly through some styled rectangles and text, and consume gigabytes of ram if you have a few tabs open. Games have directional sound effects and soundtracks. Don't need 800 Spotify engineers to pull that off. Multiplayer games solve crazy distributed system problems, making it feel like 'now' when players shoot each other, even with historical latencies of 100-200ms. AI (in terms of LLMs) seems to be a continuation of that. You used to be able to play 7 AIs on 1998 hardware, at a distinctly "non-beginner level".
- TeMPOraL 13d agoGame developers cheat like there's no tomorrow, though. In the videogame 3D graphics space, the old mantra was, "if it looks right, it's right". That barrel you shoot, is really half a barrel when you're up close, a flat rectangle when you're far, a point-with-mass + a vector for purposes of physics, and not even there for purposes of AI because pathfinding uses a precomputed graph of nodes that's carefully aligned with the map so you don't notice the enemies can noclip through everything other than floors and walls. Etc. And yes, many games would have scripted enemies or other events come out at you so you don't linger in particular areas too long, lest you spot some of the shortcuts they made. I grew up wanting to make games, spent my teenage years in hobbyist gamedev communities, and to date, this remains to me the most enjoyable and pure form of exercising software development skills.
- cm2012 13d agoIt's very different to have a generalist AI be able to win rather one trained for starcraft
- alembic_fumes 13d agoThis is a very interesting benchmark, and I think it has a lot of potential to make the speed of a model quantifiable. I'm often asking myself is it better to use higher or lower effort levels, or to maybe drop down to a "dumber" but faster model. And so using a real-time based competition as a benchmark could shed some light on this, I think. In this vein, here are what I would love to see added in this benchmark: - Include Google's Gemini models. I keep hearing Gemini being praised for its speed, and I would like to see whether that gives it a big enough edge over the bigger but slower models. - How does a Cerebras-accelerated open source model fare against a much larger but much slower frontier model? I also feel like in general there is a lot of very low-hanging fruit to start benchmarking models across the spectrum of real-time vs batch-style workloads. Perhaps Brood War sits somewhere quite near the "real-time" end of the spectrum, but what about something like a game of speed chess, or a turn-based game with time limits? I think what I would like to see the most is for someone to come up with a benchmark that supports tuning the "real-timeliness" of the benchmark, and then running a sweep of a model across the whole spectrum. That could get result in real nice graphs with multiple models on the pareto-frontier, varying based on the hosting provider and the model dimensions.
- tianqi 13d ago“Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. ” That’s me. I’ve always struggled with real-time games because I need to pause and think. While I excel at chess and board games, I’m just no good at real-time ones. At last I can only manage by sticking to a fixed set of tactics for a game, which minimizes the need for on-the-fly thinking. Seeing current models face the same difficulty leaves me with mixed feelings.
- 27388383 13d agothanks for the blog
- herodoturtle 13d agoIt was the late 90s, my very first day at a new school, I was asked to introduce myself at the front of the class, I mentioned I like computers, one guy at the back of the class blurts out “En Taro Adun” and without skipping a beat I replied “J’tokoh zohl”, and we instantly became best friends.
- egeozcan 13d agoMake the OpenAI models develop AI-scripts to play the game instead (like the good old AI-scripts, not AI as in LLMs). They are amazing at that.
- kennywinker 13d agoSo they aren’t intelligent. Like if a model can’t handle a task it hasn’t been trained on extensively, that’s not intelligence it’s memorization.
- egeozcan 13d agoNo they can handle it, but they need too much time. If you pause the game when the LLM thinks, it works. Wrong tool for the job in its current state.
- steve_taylor 13d agoFable's effort level is a significant omission, given it was ranked 3rd behind Astra xhigh and Astra medium.
- DeepYogurt 13d agoHmmmmm, almost like these models aren't generally intelligent
- rob313 13d ago"another match is already active" Really looking forward to playing- are you all limiting boxes?
- leobuskin 13d agoAstra had an explicit medium/xhigh levels, Fable - just Fable. What reasoning level was used? Why not multiple were tested? I’ve scrolled the article, but haven’t noticed any remarks about Fable’s levels.
- monk_grilla 13d agoI was looking for this comment. Why were only the Codex models exercised at each effort level?
- nrightnour 13d agotypesafe.ai would eat them all for lunch.