4 ms·
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
by Tsarp 12d ago
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
- rvz 12d ago[flagged]
- jcims 12d agoWe're allowed to have our ceremonies.
- kridsdale3 12d agoThank you. If this whole thing isn't fun, it isn't worth doing.
- user43928 12d agoYou don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
- TylerE 12d agoAbsolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
- lumirth 12d agoHave you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
- user43928 12d agoI disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting. That it isn't the most efficient way to achieve the same end result is irrelevant.
- forgot-my-pw 12d agoIt might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=codex-deepseek-v4-pro-0813-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cantigravity-sdk-gemini-3-8-flash-high%2Cmuse-code-muse-spark-1-3-max%2Ckimi-code-cli-kimi-k3%2Cclaude-code-fable-5-1-max-with-fallback%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-7-xhigh%2Cgrok-build-grok-4-6-xhigh#coding-agents-token-usage-chart-tabs https://artificialanalysis.ai/agents/coding-agents?agents=co...