4 ms·
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1
by bertili 1mo ago
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap!
Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
- cbg0 1mo agoBut is the score really reflective of the quality or are both models benchmaxxing?
- bermudi 1mo agoMuse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD. This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
- gpt5 1mo agoBoth versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
- dominotw 1mo agohow much of it is from reallocation of staff to ai training and labeling
- WASDx 1mo agoWith the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
- dakolli 1mo agoand they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
- skybrian 1mo agoI’m retired so it won’t be replacing my labor :)
- comicjk 1mo agoMy labor makes other people's lives better, so I would expect something that replaces my labor to do the same.
- dakolli 1mo agohttps://en.wikipedia.org/wiki/Commodity_fetishism https://en.wikipedia.org/wiki/Commodity_fetishism
- atemerev 1mo ago[flagged]
- MadameMinty 1mo agoBuddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.
- atemerev 1mo agoWhy "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.
- notatoad 1mo agowhen are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
- caconym_ 1mo ago+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem to inhabit.)
- jdm2212 1mo agoAnd Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
- albrewer 1mo agoA series of hot takes: Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
- jdm2212 1mo agoI think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
- zackify 1mo agoI have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
- israrkhan 1mo agoGemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens. Compare that to Muse spark 1.3 $1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing) It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.