4 ms·
Open-weight 27B hits 38% on Terminal-Bench 2.0 (Opus 4.1 hit 38% in Aug 2025)
- ubermon 5mo ago[dead]
- debpack 5mo agothis is super sick man
- ubermon 5mo agoThanks!
- timothyshen123 5mo agoInteresting find on this. Thanks for sharing
- ubermon 5mo agoThank you! I think there is a lot to dive deep later with different hardware, inference engine, prompt/harness setup etc.
- annjose 5mo ago> today's best runnable-offline model is roughly 6–8 months behind today's frontier. But it doesn't matter because frontier models were extremely good 8 months ago and we were doing real work with them. Now we have more capable open-source agents like pi and OpenCode which work well with these models. More importantly, offline models is the best choice for privacy, on-device inference and no token/cost anxiety.
- merkleforest 5mo ago> 2. How local use feels in practice Do we have stats on how does the models do on Mac M-series chips?
- ubermon 5mo agoNot yet, will conduct a more comprehensive one later