5 ms·
I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever
by Johnny_Bonk 2mo ago
I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
- dhorthy 2mo agoyeah this was just a start - the fastest cheapest thing we could try for a brand new model. I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
- Johnny_Bonk 2mo agoyeah i agree
- dan_gee 2mo ago[flagged]
- scrollaway 2mo agoHow is this useful or insightful? You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
- dan_gee 2mo agoMy point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
- Johnny_Bonk 2mo agolike all good science, measure everything
- joatmon-snoo 2mo agoThe evidence that proves a model is actually a step function change is these benchmarks. If a model isn’t a step function change? Welcome to research.
- deleted 2mo ago[deleted]