3 ms·
Yeah, I would like to know this. From my perspective, even frontier models by the big players (4o, 3.5 Sonnet) can be unreliable at times, and are at best just
by thegeomaster 2y ago
Yeah, I would like to know this. From my perspective, even frontier models by the big players (4o, 3.5 Sonnet) can be unreliable at times, and are at best just walking the line of usefulness for a lot of "exact" tasks (for me: programming, approximation and back-of-the-envelope calculations, expertise on subjects I'm unfamiliar with, a better Google, etc.).
The only deal that would make sense for me is to get something more accurate, and these open models just go in the wrong direction. I've observed similar behavior to what you mention. I'd really like to know how people use them and for what tasks so that their performance is acceptable.
In agentic settings, cost and latency are also a large factor since tokens are consumed invisibly, so I think a lot of these systems are waiting for a trifecta of better accuracy, better cost and better latency to make them viable. It's unclear that this is coming, at least it hasn't been the trend so far.
- weird-eye-issue 2y agoo1 is 10x better than 4o and 3.5 Sonnet for non-trivial coding tasks. I don't even bother with the other models for coding-related tasks, it really is a big difference.
- brookst 2y agoYep. Calling 4o a “frontier model” when o1 is available seems questionable. I only use 4o when I need web search incorporated in results.
- thegeomaster 2y agoFor isolated coding tasks I've been using o1-preview instead of Sonnet for a while now, I just didn't mention it. Haven't had a chance to test o1 proper, but I assume it's also a jump in performance. However, for more "holistic" tasks which need to take into account a larger view of some other modules/systems/interfaces, I've found o1-preview can get really confident about weirdly incorrect things that end up being harder to debug than the more straightforward hallucinations of Sonnet, and so I mostly revert to Sonnet in those cases. I tried not to make too big a fuss about the exact models I'm mentioning, since it's pretty clear that the strongest open model, Llama (discussed here), is not comparable to inference-time compute models. And for agentic settings which I mentioned above, o1 just tips way too much into expensive & slow territory to make it useful, so my prediction is that it will be limited for direct (chat-based) consumer use for the time being.