3 ms·
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or th
by caconym_ 26d ago
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
- jdm2212 26d agoAnd Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
- albrewer 25d agoA series of hot takes: Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
- jdm2212 25d agoI think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
- zackify 26d agoI have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
- KptMarchewa 25d agoNeither flash or regular glm 5.3 are close in my experience. I still prefer Sol though.
- zackify 24d agoWhat type of setup do you use? I have very small 4-5k initial context and do one task then clear. I rarely go above 150k context for most things.