4 ms·
Absolutely decimated on metrics by o4-mini, straight out of the gate, and not even that much cheaper on output tokens (o4-mini's thinking can't be turned off II
by ein0p 1y ago
Absolutely decimated on metrics by o4-mini, straight out of the gate, and not even that much cheaper on output tokens (o4-mini's thinking can't be turned off IIRC).
- gundmc 1y agoIt's good to see some actual competition on this price range! A lot of Flash 2.5's edge will depend on how well the dynamic reasoning works. It's also helpful to have _significantly_ lower input token cost for a large context use cases.
- rfw300 1y agoo4-mini does look to be a better model, but this is actually a lot cheaper! It's ~7x cheaper for both input and output tokens.
- vessenes 1y agoo4-mini costs 8x as much as 2.5 flash. I believe its useful context window is also shorter, although I haven't verified this directly.
- mccraveiro 1y ago2.5 flash with reasoning is just 20% cheaper than o4-mini
- vessenes 1y agoGood point: reasoning costs more. Also impossible to tell without tests is how verbose the reasoning mode is
- mupuff1234 1y agoNot sure "decimated" is a fitting word for "slightly higher performance on some benchmarks".
- kfajdsl 1y agoAnecdotally o4-mini doesn’t perform as well on video understanding tasks in our pipeline, and also in Cursor it seems really not great. During one session, it read the same file (same lines) several times, ran ‘python -c ‘print(“skip!”)’’ for no reason, and then got into another file reading loop. Then after asking a hypothetical about the potential performance implications of different ffmpeg flags, it claimed that it ran a test and determined conclusively that one particular set was faster, even though it hadn’t even attempted a tool call, let alone have the results from a test that didn’t exist.