2 ms·
sounds like you might need to beef up your harness first, and run multi-agent verification loops
by ernsheong 19d ago
sounds like you might need to beef up your harness first, and run multi-agent verification loops
- waldrews 19d agoThat's fine if you're doing interactive dev tasks, but we're in the large volume, cost effective, big inputs, business still with low error tolerance business, and tuned the heck out of what we can get with minimal fix cycles. Millions of cases at hundreds of thousands tokens each - after all the prefiltering by cheaper models - and the tasks still need them to do convoluted reasoning. So 'usually get it right the first time' is a big part of the cost equation.