3 ms·
Here is GLM 5.2 (https://codeinput.com/s/HAO0qTxw2ia https://codeinput.com/s/HAO0qTxw2ia) which is still inferior to Opus. I can't find Qwen 3.8 which now is my
by csomar 2mo ago
Here is GLM 5.2 (https://codeinput.com/s/HAO0qTxw2ia https://codeinput.com/s/HAO0qTxw2ia) which is still inferior to Opus. I can't find Qwen 3.8 which now is my daily driver replacing GLM. This SVG test matches my experience when working with the different models. The other models can get the details right but their output is structured in a way that makes little or no sense.
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.