4 ms·
Even o1-preview hallunicates more than e.g. Claude Sonnet 3.5 before it and only slightly better than GPT-4o, according to the paper on OpenAI's own SimpleQA be
by jug 2y ago
Even o1-preview hallunicates more than e.g. Claude Sonnet 3.5 before it and only slightly better than GPT-4o, according to the paper on OpenAI's own SimpleQA benchmark. This is precisely the problem o1 tried to tackle at great effort i.e. despite consuming far more resources as it tries to reason. So while o1 was an improvement in general quality, it is also a failure in terms of what OpenAI really wanted to see and it remains unclear what kind of future their reasoning models have. There's an emergent picture being painted of OpenAI that I'm not sure all investors are even on board on and seeing yet.
- anshumankmr 2y agoYeah even Open AI said o1-preview was not going to be a replacement for 4o. It is still exciting to see some breakthroughs in this field, but it isn't going to be like the early days of GPT-4 release.
- tiahura 2y agoYou can’t rule out a rabbit out of the hat for 5, but it sure seems like 4.0 was an inflection point and a good time to sell-out.