5 ms·
Seems like there's no official blog post with benchmark results yet. But I'm once again thankful for the Chinese AI labs for being open with their work and cont
by Reubend 4mo ago
Seems like there's no official blog post with benchmark results yet. But I'm once again thankful for the Chinese AI labs for being open with their work and contributing it to the world under permissive licenses like this. The Fable 5 fiasco is just another reminder of how valuable these things are to have.
- LaurensBER 4mo agoBased on my first impressions it's about 6 months behind the frontier labs. So very similar to Opus in January. That is, pretty damn impressive and very useable. When it comes to architecture or complex problems it does noticeable worse but I don't think anyone expected anything else. One particular interesting strong point seems to be design and user interfaces. It does seem to punch above it's weight there but that might just be personal preference.
- Lord-Jobo 4mo agoIt’s insanely impressive and I’m so glad that the space has actual competition
- becomevocal 4mo agoAppreciate the quick take! Sounds like a keeper to me. I think the Opus and Fable design (that I saw for a short while) have gotten stale
- GCUMstlyHarmls 4mo ago> I think the Opus and Fable design (that I saw for a short while) have gotten stale Can you expand on what you mean by stale? I don't get how an artefact-producer can get "stale" besides literally out-of-data information which I dont think you mean because you mention fable.
- collingreen 4mo agoI think they mean the style these tend to put out is becoming noticeable in too many places and therefore the resulting frontends feel stale, ie not "fresh" or unique
- deleted 4mo ago[deleted]
- byw 4mo ago> Opus in January So pre-nerf Opus?
- ifwinterco 4mo agoWas going to say, I don't think Opus has really got much better in the last 6mo. It just goes in cycles of being better and then being worse again, presumably based on how much Anthropic are having to optimise inference
- deleted 4mo ago[deleted]
- pastel8739 4mo agoOpus in January was right about when AI became actually useful for coding for me. So if that’s the case, that is absolutely great.
- ignoramous 4mo ago> Based on my first impressions it's about 6 months behind the frontier labs. So very similar to Opus in January. According to this one benchmark, I find it amusing that Qwen3.6 27B beats ALL "frontier lab" models on coding Kotlin: https://archive.vn/RYBCL https://archive.vn/RYBCL / https://gertlabs.com/rankings?mode=agentic_coding&language=kotlin https://gertlabs.com/rankings?mode=agentic_coding&language=k...
- ThouYS 4mo ago3.6 is an absolute beast! makes you wonder why the big heavy models are even needed?!
- jstummbillig 4mo ago> When it comes to architecture or complex problems it does noticeable worse but I don't think anyone expected anything else. So it's not really similar to opus in January?
- Eridrus 4mo agoReleasing a model without benchmarks seems to say the model is probably bad...
- deleted 4mo ago[deleted]
- vidarh 4mo agoI just ran a report from a project I'm working on that uses a mix of models, and GLM 5.1 trumped Sonnet over the last week, so I'm excited to now turn on 5.2. This is based on completion only - not quality, but that includes passing a huge test suite, and Sonnets failure rate was surprisingly bad... What I've seen from 5.1 for things like planning has certainly not read as impressive as Opus, and often even as Sonnet, but it's been a strong and steady work-horse that's just kept on actually delivering progress.
- khalic 4mo agoIt's also a reminder that as soon as Chinese models take the lead, they will switch to closed source too... so let's not be complacent, we need stronger, completely open data models, open source code, etc. to mitigate this risk
- cududa 4mo agoHow do you figure that? “also a reminder that as soon as Chinese models take the lead, they will switch to closed source too” What specifically about their release strategy “reminded” you of that conjecture? The premise that they only open source the models … because it somehow helps them leapfrog American labs, and once they actually can leapfrog them, they’d close source them, doesn’t really track for me. Am I missing something? I mean I think we need our own domestic open weight labs. I just don’t particularly understand the point you’re making
- khalic 4mo agoThe point I’m making is that this has become a strategic resource. The Chinese government allows wide sharing of their models because is weakens the US position. If Chinese models become better than Americans, do you believe the CCP will allow the free distribution of their flagship models? Think again if it’s the case.
- LogicFailsMe 4mo agoThey would still be at a significant compute disadvantage and deploying them worldwide seems to be how they work around that currently as they put together a homegrown alternative.
- khalic 4mo agoOh i don't expect this to happen any time soon, but they are making progress on the UV lithography side, so it's just a matter of time until it becomes a TW race, and they have the advantage on that terrain.