3 ms·
For similarly sized models, not looking very good on the slightly-less-benchmaxxed Terminal-Bench 2.0: Laguna XS.2 33B-A3B params: 30.6 Qwen 3.6 35B-A
by jaen 6mo ago
For similarly sized models, not looking very good on the slightly-less-benchmaxxed Terminal-Bench 2.0:
Laguna XS.2 33B-A3B params: 30.6
Qwen 3.6 35B-A3B : 51.5
Devstral 2 123B : 31.2
Quite a huge lead for Qwen... well, at least it's catching up to other smaller Western labs.
- megavon 6mo agoNeed to look at SWEBench-Pro, it's super competitive. Suspect they'll catch up given the longer-tail on TB scores.
- jaen 6mo agoJust by the (lack of) inter-model variance, I don't think SWEBench-Pro does a very good job of representing model capability. Terminal-Bench seems more challenging and separates the wheat from the chaff. Also, *ops work, which in my experience can actually be more complicated than SWE is underrepresented there obviously.