3 ms·
For isolated coding tasks I've been using o1-preview instead of Sonnet for a while now, I just didn't mention it. Haven't had a chance to test o1 proper, but I
by thegeomaster 2y ago
For isolated coding tasks I've been using o1-preview instead of Sonnet for a while now, I just didn't mention it. Haven't had a chance to test o1 proper, but I assume it's also a jump in performance. However, for more "holistic" tasks which need to take into account a larger view of some other modules/systems/interfaces, I've found o1-preview can get really confident about weirdly incorrect things that end up being harder to debug than the more straightforward hallucinations of Sonnet, and so I mostly revert to Sonnet in those cases.
I tried not to make too big a fuss about the exact models I'm mentioning, since it's pretty clear that the strongest open model, Llama (discussed here), is not comparable to inference-time compute models.
And for agentic settings which I mentioned above, o1 just tips way too much into expensive & slow territory to make it useful, so my prediction is that it will be limited for direct (chat-based) consumer use for the time being.