4 ms·
https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F58dc4e8ec482330856fce89dac670727 https://tools.simonwillison.
by simonw 5d ago
https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F58dc4e8ec482330856fce89dac670727 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8ab126bda2b384264b3ad931e3ebb8b4 https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0 https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
- MattDamonSpace 5d agoAre there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
- datsci_est_2015 5d agoPoor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
- kiliancs 5d agoWhat is the default reasoning level?
- forgot-my-pw 5d agoI tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=codex-deepseek-v4-pro-0813-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cantigravity-sdk-gemini-3-8-flash-high%2Cmuse-code-muse-spark-1-3-max%2Ckimi-code-cli-kimi-k3%2Cclaude-code-fable-5-1-max-with-fallback%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-7-xhigh%2Cgrok-build-grok-4-6-xhigh#coding-agents-token-usage-chart-tabs https://artificialanalysis.ai/agents/coding-agents?agents=co...
- TomGarden 5d agoI think these are the worst I've seen, at least in some time. It's a silly benchmark though, not sure what to make of it
- daveguy 5d agoHahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?