7 ms·
Agent swarms and the new model economics
- dctwin 2mo agoAm I reading this right? Opus + Composer did a comparable job to Fable, at ~1/19th the price, and half the LoC?
- LostMyLogin 2mo agoI'm struggling to use Fable correctly. Everything I do ends up getting flagged and reverted back to Opus. Need to do some digging.
- Footprint0521 2mo agoAnthropic wants it that way so they can’t be sued… the correct answer is don’t, switch to Kimi K3 for much more usage, the same quality of model, and no hidden practices :/
- KyleTheDev 2mo agoYea, looks like it (at least for their v2). Fable did much better in V1, but once they added their tooling around it, Opus + Composer ended up doing better (>0.8 Grade on the SQLite chart) for the final product. Granted, Fable reached a 'good' result (0.7~) in a much faster timeframe.
- espetro 2mo agoPoint is, frontier models don't necessarily are better than a swarm of non-frontier ones for these scoped problems. IMO frotnier models excel at underspec'd or more complex problems where ideation and exploration are key (and that's where they become crazy expensive).
- htrp 3mo ago>The browser swarm from earlier this year peaked at roughly 1,000 commits per hour on Git. The new system peaks at around 1,000 commits per second. >To facilitate this rate of activity, we built a new version control system (VCS) from scratch. Throughput was not the only reason to own this layer. Every change in the system passes through the VCS, so it is where collisions first become visible, and several of the coordination mechanisms in the next section are implemented directly inside of it. Talk about inventing the universe to make a button.
- jgalt212 3mo agoIt's hard not read such quotes and immediately think of the Infinite Monkey Theorem. https://en.wikipedia.org/wiki/Infinite_monkey_theorem https://en.wikipedia.org/wiki/Infinite_monkey_theorem
- whatever1 3mo agoIf at the end Shakespeare-level literature is produced, does it matter whether we arrived there by random keystrokes?
- voidhorse 3mo agoYes because the keystrokes aren't free and the world is finite. We need to pay these monkeys in bananas.
- qjack 3mo agoThe problem is the infinite other literature produced along the way. It doesn't work to throw more monkeys at the problem of verification, and checking for matches against existing Shakespeare plays is cheating.
- subygan 2mo agodon't worry, we can get another swarm of random monkeys typing and decide what ones to eliminate . tijebs gi yo,
- deleted 2mo ago[deleted]
- confidantlake 2mo agoYes because it is akin to producing a universe of slop and trying to find the single atom of gold.
- w-ll 2mo agoShakespeare-level literature was produced. We are the infinite monkeies, and one was shakespeare.
- handfuloflight 3mo ago> To test that progress, we returned to a task the old swarm had struggled with: building SQLite from scratch, in Rust, from nothing but its documentation. Isn't SQLite's source code in its training data?
- Romario77 3mo agonot in Rust. might be different enough ...
- hiddendoom45 3mo agoThere's Turso which is a sqlite rewrite in Rust with some additional features. It's open source so I'd imagine it would be in the training data.
- kgeist 3mo agoEven if no Rust code for it was seen during training, an LLM can trivially transpile SQLite's C codebase to Rust on the fly. For example, I just asked ChatGPT to write John Carmack's famous Fast Inverse Square Root algorithm in Erlang, without searching online or thinking, and it transpiled it immediately (while also extracting the knowledge in the same step). SQLite's semantics/code are stored in the middle layers of an LLM, and the last layers are able to convert it into any representation, as conditioned by the prompt. Cursor's experiment is deeply flawed because they merely extracted the model's compressed, lossy knowledge of SQLite's codebase and then just ran a bunch of tests/fixing rounds to make up for the lossiness. The claim that the agents built it from scratch is false.
- wrs 3mo agoThen you would expect the implementation to be structured the same as SQLite, and having glanced at the result, it looks like at least some things aren't. For example, it seems to use an operator-tree executor rather than SQLite's bytecode interpreter.
- deleted 2mo ago
- anthonypasq 3mo agoLove to see these crazy kinds of experiments going on. Even if this doesn't 100% work or is prohibitively expensive for now, these are glimpses into the future in the same way people were talking about coding agents in 2023 when we just had tab complete.
- sawyers 2mo agoIt shouldn't be prohibitively expensive because the more you break the work into smaller parts, theoretically the smaller a model can be that completes it. Therefore your agent swarms don't have to be cutting edge API calls, they can be hosted on clusters of tiny cheap computers with small CPU GPU. For orgs you could basically invest in a local server that has 100s of small computers. You can use engineers and frontier models to plan and do task subdivision. And then shit that out to the local cluster that's basically handling each small task without knowing what it's supposed to do. The real issue is the middle. Testing, verifying interconnectivity of mid level abstractions, building that amount of tech debt at such a high rate and actually being aware in any way what it is that you've built. I think for most MBAs, considering the US and a significant amount of global markets only care about short term market upside, and AI is still in the middle of the hype cycle, that's not a serious issue for anyone who is profit oriented. It's only a real issue for the losers like you and me who care about sustainability and infrastructure that survives longer than one market hurricane season.
- Schlagbohrer 2mo agoThe twin trends of AI slop being churned out at tremendous rate and deployed in everything, plus frontier AIs becoming really good at exploiting many small security holes to pwn boxen, is going to collide in an amazing cyberstorm that was predicted by every cyberpunk novel from the 80s X-D and we are gonna have front row seats 8-D
- cookiengineer 2mo agoThis so much! For my current setup, the most efficient way is to use larger models for coordination, but a heretic'ed qwen 30B model for the implementations. If you build your agentic environment around specifications and test coverage tied to symbols, you can do a lot of parallelization of agent work. If you then separate the filesystem and tool read/write access by agent roles and policies (e.g. coder not allowed to modify unit tests, tester not allowed to modify code files) then you can force them to use a centralized per symbol/per contract specification tool. And then you can just let agents discuss issues, where the messaging threads are also tied to the same symbols. [1] Exocomp, highly experimental: https://github.com/cookiengineer/exocomp https://github.com/cookiengineer/exocomp
- mccoyb 3mo agoI find these blog posts (and the originals, with Anthropic's C compiler and Cursor's browser) somewhat funny, as if they have this enormous power to build ... but they can't build something unique or new. Like the software sucks, but look how powerful the process is (the models are indeed powerful). And it's a bit of a shame: by virtue of their position (their embedding in the fabric of venture capitalism), it seems like they can only make a subset of things -- what they can make is dictated enormously by capital, as they are engines of capital. Not sure the point I'm trying to make, I just find it amusing. Perhaps the point is that it might be more worthwhile to give independent creators a billion dollars to play around with agent swarms if we want to keep diversity in the evolutionary algorithm that is the software industry high.
- deleted 3mo ago[deleted]
- brap 3mo agoBecause it's mostly a load of crap. Most of these reusable, automated workflows sound really cool in theory, but in practice either they require a lot of customization (not really reusable) or a lot of babysitting (not really automated). AI has improved my productivity significantly, I'm not a hater, but we also have to be honest about its (current) limitations. The whole "software factory" thing is just a pipe dream people sell to execs / investors / online courses.
- thewhitetulip 2mo agoYes and this can be proven because what Anthropic says and does is vastly different They say SaaS is dead and you can vibe code any SaaS using Opus and yet despite having Mythos with them they have to use JS terminal to run Claude Code. Why don't they just vibe code CC? Why doesn't Dario have a 100000 agent swarm to run entirety of Anthropic or at least software department?
- jnwatson 2mo agoThey've said for some time that Claude Code is 90% LLM generated code.
- kimonsodu 3mo ago[flagged]
- whinvik 3mo agoI would have loved to see more of the harness engineering shared as code. Instead we are left with only the outcome. I guess that makes sense since the harness is the product in the case of Cursor.
- dakolli 3mo agoI call this junk "meta-agentic engineering", it reminds me of people who have the coolest nvim configs, spend hundreds of hours customizing it but ultimately get less work done than the guy with minimal workflows, if any at all. I look on twitter and it's just people building tools for agents to use agents, some weird customization loop going on in the LLM space right now. Ultimately these are trends pushed on us from model providers because they 10x token consumption. Its literally just BS trends to increase revenue at these companies, most of it is largely useless.
- svachalek 3mo agoSome of this stuff is ludicrous. I finally tried /loop last week and discovered every loop iteration passes the entire context history to the model. So pretty quickly you're running a full 1M context window, without cache, likely just to check if something is ready or needs to be done. It's miserably terrible engineering unless your entire and only goal is to burn tokens.
- Wowfunhappy 3mo agoWait, why isn't context cached with /loop?
- Olscore 3mo ago[dead]
- deleted 3mo ago[deleted]
- shay_ker 3mo agoHow do we know if these models weren’t trained on Turso’s rewrite of SQLite in Rust? It seems both likely that they were and impossible to remove that code from pretraining. Doesn’t that make this just about LLM memorization of the training set? What am I missing?
- handfuloflight 3mo agoThis is more about automated long horizon work.
- adamtaylor_13 3mo agoWouldn't that be in the authors' interest to disclose, if true?
- IanCal 3mo ago> What am I missing? They’re testing the same models on the same task but different ways of organising the swarms and the new approach works better.
- DekryptLabs 3mo ago[dead]
- pianopatrick 3mo agoI wonder what would happen if the agents did not have the spec. Like if you just said "build a simple SQL database that works as a single file", what would the agent swarm come up with.
- Elad-Rez 3mo ago[flagged]
- bofadeez 3mo agoThe issue is that the only models that could be trusted to work autonomously cost more than a human employee
- uf00lme 2mo agoCursor is saying we need to get improve the harness and tooling, but the models are getting better. One would have to assume that will continue and that cheaper/smaller models will continue to become even more useful.
- onion2k 2mo agoIf there's a human who could write SQlite in an hour we should be paying them at least that much.
- vessenes 3mo agoThis is super fascinating, and I loved seeing testing of where exactly you need frontier intelligence -- looks like coordination / planning, but not coding right now -- the article's a bit of a tease, as we can't play with such a harness, or their new version control system or get a workable artifact out of it. That said, I love the work on figuring out these harness coordination jobs. While there are analogs to human management there is also this enticing feeling that, since the models are broadly deterministic, we might be able to get repeatable science-type lessons about managing them with enough testing.
- doctoboggan 3mo agoI feel bad for the engineers at Cursor who have to use Grok in these sorts of experiments.
- ryankuykendall 2mo agoThey received $60 billion to use grok and make it better. I think they going to be okay.
- dougSF70 2mo agoCan a swarm of agents write the 835 page manual?
- overgard 2mo agoGood news, the new manual will be 83500 pages.
- lubujackson 2mo agoThat's 100x productivity right there.
- dymk 2mo agoThis post describes ideas found in existing agent orchestrator systems since they've existed. You'll find the same "top level agent that breaks a problem down into spec / plan / implement / verify" pattern in get-shit-done or superpowers. Is the original part the VCS that works at 1000 commits/second? Or spending a bunch of money comparing models?
- smoyer 2mo agoThis is almost a year behind Steve Yegge's first post on beads. Gas Town and Gas City provide orchestration for the swarm. So far I haven't seen a perfect implementation but this idea isn't new.
- nullsanity 2mo agoDoes every peice of prose published here have to be novel? is it not enough for you to read a new. perspective?
- sudb 2mo agoI don't think Gas Town was actually all that much ahead of the curve - see Cursor's swarm post from January this year, mentioned at the very top of this blog post. I also don't know that beads is a very strong example that Yegge knows what he's doing - it's famously been described as (pseudo)malware.
- IceDane 2mo agoEverything Steve yegge has done has been trash. That's why nobody is talking about beads or gas town. It was clear even in the beginning that it was a borderline AI-psychosis-fueled trash fire.
- anentropic 2mo agoAt the same time a lot of the stuff in this Cursor post sounds like things that Gastown was doing, albeit more sober and thought through
- smoyer 2mo agoAgreed ... I'm not talking about the implementation but rather the idea that this research is novel.
- IceDane 2mo agoI think if you're suggesting that Steve Yegge somehow invented having agents collaborate, you're way off base. People started talking about agents collaborating very, very early on, and people have been doing experiments like this for ages. There have been libraries to build e.g. graph-based (possibly multi-agent) workflows from the early days, before even structured output was a standard thing in the APIs. The only thing Yegge did was come up with really stupid, convoluted and anthropomorphized language to describe the process and then write unhinged articles about it
- timcobb 2mo ago> One reading is that it was more productive. Another is that most of those commits were busywork (thrash, contention, churn). Kudos
- alienbaby 2mo agoTinfoil hat: what if the prices are artificially inflated deliberately to price out casual users being able to field large swarms of frontier model agents because it could be just too dangerous.
- onion2k 2mo agoThe fact that the swarm needs high quality docs to work from will keep most of the tech industry safe.
- trevswarm 2mo agoEven at a much smaller scale of a few dozen parallel agents working I've found real benefit to a structured hierarchy of agents. Not just keeping the context clean and focused at the right level, but if you find that something has just gone off the rails, you can more easily rip out or even just fully delete a leg of work without damaging the original design. It's also nice when you find a design change is needed and you can communicate that to the top level and have it efficiently propagate down where it matters. I like it so much I even recorded a video walking through the process! https://www.youtube.com/watch?v=efUmcCiRoDU https://www.youtube.com/watch?v=efUmcCiRoDU (I do walk through how I do this in our app DevSwarm here)
- agustechbro 2mo agoWow, people have been working hard for a long time in Turso: https://turso.tech/blog/introducing-limbo-a-complete-rewrite-of-sqlite-in-rust https://turso.tech/blog/introducing-limbo-a-complete-rewrite... to do the same cursor did in justo no time?
- tal-onn 2mo ago[flagged]
- edg5000 2mo agoI initially thought that getting agents to work for longer and in large groups was the future, but I'm increasingly thinking that, at least for engineering, just one thread makes more sense. The agent pulls things into context as needed. One thing that I've been experimenting with is also letting the agent remove things from context, such as files. But just adding to the context and compacting when it's full seems like it might beat a lot of more advanced options. Because the model is good, it knows what to put in the summary; just enough for the model to be able to rebuild the context from that seed (e.g. pulling in relevant files/data into the context).
- Schlagbohrer 2mo agoHave you used this method yourself for long workflows and with contexts approaching 1 million tokens? It doesn't work very well. LLM context is nearly half unusable.
- noosphr 2mo agoI've found that model performance starts substantially degrading after 10%.
- Schlagbohrer 2mo agoI see with qwen3.6-35B running in Pi that it gets stuck in loops when the 262k context gets more than 30% full.
- edg5000 2mo agoI always let GPT 5.5/5.6 sol reach compatction. I'm only offered a 258K window, but I think once they release 1M to normal plan users, it will be usable across the whole 1M. At least with Opus I could use the whole 1M. I'd argue that the more context I use, the better performance gets. It has more info already available. I don't notice degradation.
- asim 2mo agoAnything architectural like this is super interesting. We keep getting more powerful models sure. We're getting agents ok. But this coordination of many agents. This is the part that's really going to scale very fast and across many many machines in parallel. Very cool.
- Schlagbohrer 2mo agoAmazing test. I wish I could spend $2000 of compute on having my own swarm build some weird piece of software with 1000 commits per second.
- lifty 2mo agoPerhaps you can get hired by a company to do that. You can't compare yourself as an individual with a company with a large budget.
- SnehRJoshi 2mo ago[flagged]
- alexmercerdev 2mo ago[dead]
- kademolu 2mo agoIf we take coding off the table i don't see the case for large agent swarms. More is not always better in my experience experimenting with agents
- s08148692 2mo agoIf you can't see cases for it you may need some glasses. Replace the graph in the docs of Top-level planner -> mid-level planners -> worker leaf nodes and you get an org chart of CEO -> executives, N-levels of management, and ICs at the leafs. It directly maps to almost any office-based company that can operate completely digitally. It's obviously in its infancy but fully agentic companies are becoming a possibility, given the right harness, models and tools
- kademolu 2mo agoLets separate "workers" from "decision seats", you might need a lot of workers which typically don't actual require inference and might actually just be cheap solvers but actually less decision seats. So i guess it depends on what you are calling agents in your scenario.
- feiz45607 2mo ago[flagged]
- notahan 2mo agoThese tests are interesting but I can't help feeling that these are benchmaxxing results. "We made our harness make sqlite in rust" sounds like a much easier probelm statement than having to recreate an application like facebook. I think the concept of having to integrate is the real challenge. Re-writing code in new languages based on decent docs is a difficult but less useful measure of how AI can help replace engineers. Imagine this: you need to build a new CRUD app but your AI agent spends 20M toks on just writing it's own custom login function, that's a waste of time and money. I would not like to have a new system developed for me in such cases. I think we need a test where the agents are more specifically tasked to find technologies worth integrating and hence write the fewest lines of code.
- lumost 2mo agoThe nature of ML is to reach for easily measured proxy problems which can then be optimized. Why worry about a backoffice integrated micro-service you have no data on when you aren't sure that you can design and implement a rewrite of an OSS project?
- paraschopra 2mo agoI think experiments like these are exciting, but they only work when you have the full suite against which you calibrate. What I'd love to see is a small group of humans steering a million agents.
- stillpointlab 2mo agoToday it is SQLite, tomorrow Linux? I'm not sure of the long-term viability of these brute force swarm approaches, but clearly the AI companies dig the marketing they get out of it. But the fact that they get so far at this stage suggests we have a very interesting 2 years coming up.
- thisismyopinion 2mo ago1000 commits of slop per second.
- dougSF70 2mo agoThey have invented software time travel: Write software => then write manual => then feed manual to cursor to write the software.
- cregy 2mo agoSeems an extreme system looking for a use case, but it's great to see research being done in this area