3 ms·
We sure did. It's a great writer, better in a harness, will process lots large context, but complex reasoning with convoluted rules and low hallucination toler
by waldrews 22d ago
We sure did. It's a great writer, better in a harness, will process lots large context, but complex reasoning with convoluted rules and low hallucination tolerance? That's still larger model territory.
- ernsheong 22d agoThere's still 3.1 Pro though, as ancient as it sounds now
- waldrews 22d agoYup. The problem is that it's bizarrely still not in General Availability status.
- ernsheong 21d agosounds like you might need to beef up your harness first, and run multi-agent verification loops
- waldrews 20d agoThat's fine if you're doing interactive dev tasks, but we're in the large volume, cost effective, big inputs, business still with low error tolerance business, and tuned the heck out of what we can get with minimal fix cycles. Millions of cases at hundreds of thousands tokens each - after all the prefiltering by cheaper models - and the tasks still need them to do convoluted reasoning. So 'usually get it right the first time' is a big part of the cost equation.