4 ms·
We separated out the runbook such that each step is a separate LLM in the agent. Between each step, there's sort of a "supervisor" that ensures that the step wa
by wilson090 2y ago
We separated out the runbook such that each step is a separate LLM in the agent. Between each step, there's sort of a "supervisor" that ensures that the step was completed correctly, and then routes to another step based on the results. So in reality, a single step failing requires two hallucinations. Hallucinations are also not a fixed percentage across all calls -- you can make them less likely by maintaining focused goals (this is why we made runbooks agentic rather than a single long conversation)
- threeseed 2y agoAnd what is your average error rate per runbook step.
- jtsaw 2y agoone thing we're experimenting to help with the hallucinations/error rate issue is using a committee framework where we take a majority vote. If the error rate of 1 expert is 5%, then for a committee of 10 experts, the probability a majority of the committee errors is around 0.00276% (binomial distribution with p=0.05). For 10 steps, this would be an error rate of 0.0276%
- threeseed 2y agoPretty bad maths there. Those committee members are not independent. They are highly correlated even amongst LLMs from different vendors.
- jtsaw 2y agoI'm not sure they are highly correlated. A committee uses the same LLM with the same input context to generate different outputs. Given the same context LLMs should produce the same next token output distribution (assuming fixed model parameters, temperature, etc). So, while tokens in a specific output are highly correlated, complete outputs should be independent since they are generated independently from the same distribution. You are right they are not iid but the calculation was just a simplification.
- nerdjon 2y agoThat second step hallucinating is far more likely when you are feeding it incorrect information from the first hallucination. LLM's are very easy to manipulate. At one point with a system prompt telling Claude it was OpenAI, I was able to ask what its model is and it would confidently tell me it was OpenAI. Garbage data in, garbage data out. Admittedly that is an extreme case, but you're giving that second prompt wrong data in the hopes that it will identify it instead of just thinking it's fine when it is part of its new context.
- jtsaw 2y agoyea. We're definitely concerned about hallucinations and are using a variety of techniques to try and mitigate it (there's some existing discussion here, but using committees and sub-agents responsible for smaller tasks has helped). What's helped the most, though, is using cluster information to back up decision making. That way we know the data it's considering isn't garbage, and the outputs are backed up by actual data.