4 ms·
Every time I read something like this, it strikes me as an attempt to convince people that various people-management memes are still going to be relevant moving
by alphazard 8mo ago
Every time I read something like this, it strikes me as an attempt to convince people that various people-management memes are still going to be relevant moving forward.
Or even that they currently work when used on humans today.
The reality is these roles don't even work in human organizations today. Classic "job_description == bottom_of_funnel_competency" fallacy.
If they make the LLMs more productive, it is probably explained by a less complicated phenomenon that has nothing to do with the names of the roles, or their descriptions.
Adversarial techniques work well for ensuring quality, parallelism is obviously useful, important decisions should be made by stronger models, and using the weakest model for the job helps keep costs down.
- rlayton2 8mo agoMy understanding is that the main reason splitting up work is effective is context management. For instance, if an agent only has to be concerned with one task, its context can be massively reduced. Further, the next agent can just be told the outcome, it also has reduced context load, because it doesn't need to do the inner workings, just know what the result is. For instance, a security testing agent just needs to review code against a set of security rules, and then list the problems. The next agent then just gets a list of problems to fix, without needing a full history of working it out.
- fphhotchips 8mo agoWhich, ultimately, is not such a big difference to the reason we split up work for humans, either. Human job specialization is just context management over the course of 30 years.
- miki123211 8mo ago> Which, ultimately, is not such a big difference to the reason we split up work for humans, That's mostly for throughput, and context management. It's context management in that no human knows everything, but that's also throughput in a way because of how human learning works.
- purplepatrick 8mo agoI’ve found that task isolation, rather than preserving your current session’s context budget, is where subagents shine. In other words, when I have a task that specifically should not have project context, then subagents are great. Claude will also summon these “swarms” for the same reason. For example, you can ask it to analyze a specific issue from multiple relevant POVs, and it will create multiple specialized agents. However, without fail, I’ve found that creating a subagent for a task that requires project context will result in worse outcomes than using “main CC”, because the sub simply doesn’t receive enough context.
- XenophileJKO 8mo agoSo two things.. Yes this helps with context and is a primary reason to break out the sub-agents. However one of the bigger things is by having a focus on a specific task or a role, you force the LLM to "pay attention" to certain aspects. The models have finite attention and if you ask them to pay attention to "all things".. they just ignore some. The act of forcing the model to pay attention can be acoomplished in alternative ways (defined process, commitee formation in single prompt, etc.), but defining personas at the sub-agent is one of the most efficient ways to encode a world view and responsibilities, vs explicitly listing them.
- lwhi 8mo agoWhat do you think context is, if not 'attention'?
- XenophileJKO 8mo agoContext is the information you give the model, attention is what parts it focuses on. And this is finite in capacity and emergent from the architecture.
- lwhi 8mo agoSo attention is based on a smaller subset of context?
- chickensong 8mo agoYou can create a context that includes info and instructions, but the agent may not pay attention to everything in the context, even if context usage is low.
- 3371 8mo agoIMO "Attention" is an abstraction over the result of prompt engineering, the chain reaction of input converging the output (both "thinking" and response).
- ttoinou 8mo agoDevelopers do want managers actually, to simplify their daily lives. Otherwise they would self manage themselves better and keep more of the share of revenues for them
- shermantanktop 8mo agoUnfortunately some managers get lonely and want a friendly face in their org meetings, or can’t answer any technical questions, or aren’t actually tracking what their team is doing. And so they pull in an engineer from their team. Being a manager is a hard job but the failure mode usually means an engineer is now doing something extra.
- simondotau 8mo agoI suppose it’s could end up being an LLM variant of Conway’s Law. “Organizations are constrained to produce designs which are copies of the communication structures of these organizations.” https://en.wikipedia.org/wiki/Conway%27s_law https://en.wikipedia.org/wiki/Conway%27s_law
- _kb 8mo agoIf so, one benefit is you can quickly and safely mix up your set of agents (a la Inverse Conway Manoeuvre) without the downsides that normally entails (people being forced to move teams or change how they work).
- miki123211 8mo agoI think it's just the opposite, as LLMs feed on human language. "You are a scrum master." Automatically encodes most of what the LLM needs to know. Trying to describe the same role in a prompt would be a lot more difficult. Maybe a different separation of roles would be more efficient in theory, but an LLM understands "you are a scrum master" from the get go, while "you are a zhydgry bhnklorts" needs explanation.
- joshuaisaact 8mo agoThis has been pretty comprehensively disproven: https://arxiv.org/abs/2311.10054 https://arxiv.org/abs/2311.10054 Key findings: -Tested 162 personas across 6 types of interpersonal relationships and 8 domains of expertise, with 4 LLM families and 2,410 factual questions -Adding personas in system prompts does not improve model performance compared to the control setting where no persona is added -Automatically identifying the best persona is challenging, with predictions often performing no better than random selection -While adding a persona may lead to performance gains in certain settings, the effect of each persona can be largely random Fun piece of trivia - the paper was originally designed to prove the opposite result (that personas make LLMs better). They revised it when they saw the data completely disproved their original hypothesis.
- jimkleiber 8mo agoHow well does such llm research hold up as new models are released?
- dexdal 8mo agoMost model research decays because the evaluation harness isn’t treated as a stable artefact. If you freeze the tasks, acceptance criteria, and measurement method, you can swap models and still compare apples to apples. Without that, each release forces a reset and people mistake novelty for progress.
- solumunus 8mo agoOne study has “comprehensively disproven” something for you? You must be getting misled left right and centre if that’s how you absorb study results.
- ljm 8mo agoIt shows me that there doesn’t appear to be an escape from Conway’s Law, even when you replace the people in an organisation with machines. Fundamentally, the problem is still being explored from the perspective of an organisation of people and it follows what we’ve experienced to work well (or as well as we can manage).
- generallyjosh 8mo agoI do think there is some actual value in telling an LLM "you are an expert code reviewer". You really do tend to get better results in the output When you think about what an LLM is, it makes more sense. It causes a strong activation for neorons related to "code review", and so the model's output sounds more like a code review.
- zhenyakovalyov 8mo agoi guess, as a human it’s easier to reason about a multi-agent system when the roles are split intuitively, as we all have mental models. but i agree - it’s a bit redundant/unnecessary