19 ms·
Building more with GPT-5.1-Codex-Max
- deleted 11mo ago[deleted]
- iamronaldo 11mo agoThat was quick
- bigyabai 11mo agoMy first thought was "they must not be seeing as many Claude Code conversions as they hoped"
- the_duke 11mo agoI bet they just wanted to counter Gemini 3 and stay on top of the leaderboards for coding, and were preparing this for a while to push out alongside Gemini 3.
- giancarlostoro 11mo agoWhenever one of them releases a milestone release the rest start publishing big milestones too. I'm waiting for Opus 5 next.
- LZ_Khan 11mo agoall i care about is performance on metr benchmark
- Reubend 11mo agoOpenAI likes to time their announcements alongside major competitor announcements to suck up some of the hype. (See for instance the announcement of GPT-4o a single day before Google's IO conference) They were probably sitting on this for a while. That makes me think this is a fairly incremental update for Codex.
- Palmik 11mo agoGPT 5.1 / Codex already beats Gemini 3 on SWE Bench Verified and Terminal Bench and this pushes the gap further. Seems like a decent improvement.
- knowriju 11mo agoWould it be fair to compare a generic model with a model finetuned for coding?
- skhameneh 11mo agoThere’s been community commentary that many of the GPT models are a tad overfitted WRT benchmarks. Benchmarks are not representative of end user experiences. That’s not to say the benchmarks aren’t useful at all, but are only useful as a subjective indicator.
- bugglebeetle 11mo agoThat’s how the game is played. We should be grateful for all the competition that is driving these improvements, not whinging about the realities of what companies have to do to contest each other’s position.
- johnecheck 11mo agoIt's funny, this release comes right after the Gemini 3 release that coincided with day 1 of Microsoft's Ignite conference.
- peab 11mo agoit's really getting old
- johnwheeler 11mo agoGemini is eating their lunch, and OpenAI knows it.
- echelon 11mo ago
- spmartin823 11mo agoI still want something no one has, which is the ability to launch agents in different git worktrees simultaneously and check the results out on my main branch for testing when they are finished.
- bradly 11mo agoWould this be similar to how Charlie and Jules work?
- cube2222 11mo agoI think I’ve described how I achieve kinda your desired workflow in a comment yesterday [0]. [0]: https://news.ycombinator.com/item?id=45970668 https://news.ycombinator.com/item?id=45970668
- agentifysh 11mo agoha! very interesting how slept on jj is its been essential to my workflow as well i use both jj and git and jj is great for just creating a snapshot that i can revert to incase it fails im still exploring it to see what else i can do with it for agentic use
- agentifysh 11mo agolots of tools that do this and I ended up going down this rabbit hole something that could just plug in to codex instead of requiring a fork http://github.com/agentify-sh/10x http://github.com/agentify-sh/10x does minimal overhead with agent orchestration (its just a bash/typescript) as its main focus was adding enhancements to codex like double redundant checkpoint via git and jj (lessons learned from codex being git reset --hard happy), something like claude skills (just a bunch of mds that steer it towards specific activity like think, plan, execute), timeout wrappers (to get you unstuck if codex waits a long time), blacklist commands during yolo (rm -rf, git reset banned even if it by small chance run it) MIT licensed you can work sequentially (subagents launch one after the other) or parallel (worktrees) but tbh sequentially is better because you understand what is going on with parallel it might be best for dealing with tests and UI.
- agentifysh 11mo agoso this was arctic fox it seems, lot of us ended up downgrading to codex 5.0 because of the token burn was too much, i see codex max is a step up which is welcome but still unsure if they solved that github issue around tool use that impacts tokens going to wait and see after being burned by 5.1 before i upgrade back to 0.58 gemini 3 has been a let down tbh to see agentic coding wasn't a top priority im sticking with codex for now and using gemini 3 for frontend
- GenerWork 11mo agoHave you found that Gemini is better than Codex for front end generation? I'm trying to bring some Figma screens into a small React project I have, and Codex will occasionally screw up the implementation despite the fact that I'm using the MCP server.
- jasonthorsness 11mo ago"Starting today, GPT‑5.1-Codex-Max will replace GPT‑5.1-Codex as the default model in Codex surfaces." Wow, I spent last weekend using a tag-team of Claude and Codex and found Codex to more often get better results (TypeScript physics/graphics application). I probably only wrote a few hundred lines of code out of many thousands; it did a really good job. Now I guess I'll ask the new Codex to review the work of the old!
- taurath 11mo agoThese 2 sentences right next to each other stood out to me: > a new step towards becoming a reliable coding partner > GPT‑5.1-Codex-Max is built for long-running, detailed work Does this not sound contradictory? It’s been the shorter form work that has built what little confidence I have in these as a coding partner - a model that goes off and does work without supervision is not a partner to me.
- causal 11mo agoAbsolutely contradictory. The long-running tendency for Codex is why I cannot understand the hype around it: if you bother to watch what it does and read its code the approaches it takes are absolutely horrifying. It would rather rewrite a TLS library from scratch than bother to ask you if the network is available.
- keeganpoppen 11mo agothese things are actually fixable with prompting. is it easy? no. is it PEBKaC if you don’t do anything to change course as it builds a TLS library? yes, but paperclip maximized! xD
- causal 11mo agoOr you can have a model with some semblance of common sense that will stop and say "Hey I can I have access to the network to do X?" Codex feels like a tool designed to run after all the humans are gone.
- meowface 11mo ago>It would rather rewrite a TLS library from scratch than bother to ask you if the network is available. This is definitely one of the biggest issues with coding agents at the moment. That said, from my experience, Codex so often does things that are so useful and save me so much time that the occasional "oh god what the hell did it just go off and do" are an acceptable cost for me. I regularly get great results with open-ended prompts and agents that spend 15+ minutes working on the task. I'm sure they'll eventually get better at common sense understanding of what kind of work is wasteful/absurd.
- simianwords 11mo ago> Compaction enables GPT‑5.1-Codex-Max to complete tasks that would have previously failed due to context-window limits, such as complex refactors and long-running agent loops by pruning its history while preserving the most important context over long horizons. In Codex applications, GPT‑5.1-Codex-Max automatically compacts its session when it approaches its context window limit, giving it a fresh context window. It repeats this process until the task is completed. Wouldn't the model automatically do that using attention techniques? Why do you need to do it at the token layer and not leave it to the model to automatically decide which tokens are worth paying attention to?
- qsort 11mo ago> due to context-window limits
- simianwords 11mo agocontext window is not some physical barrier but rather the attention just getting saturated. what did i get wrong here?
- qsort 11mo ago> what did i get wrong here? You don't know how an LLM works and you are operating on flawed anthropomorphic metaphors. Ask a frontier LLM what a context window is, it will tell you.
- deleted 11mo ago[deleted]
- Palmik 11mo agoIt's a fair question, even if it might be coming from a place of misunderstanding. For example, DeepSeek 3.2, which employs sparse attention [1], is not only faster with long context than normal 3.1, but also seems to be better (perhaps thanks to reducing the noise?). [1] It uses still quadratic router, but it's small, so it scales well in practice. https://api-docs.deepseek.com/news/news250929 https://api-docs.deepseek.com/news/news250929
- hansonw 11mo agoRest assured that we are better at training models than naming them ;D - New benchmark SOTAs with 77.9% on SWE-Bench-Verified, 79.9% on SWE-Lancer, and 58.1% on TerminalBench 2.0 - Natively trained to work across many hours across multiple context windows via compaction - 30% more token-efficient at the same reasoning level across many tasks Let us know what you think!
- agentifysh 11mo agodid you address this https://github.com/openai/codex/issues/6426 https://github.com/openai/codex/issues/6426 ? how much more token efficient is this compared to 5.0 had to use 5.0 because 5.1 was eating tokens like crazy and seemed like a slight incremental improvement barely noticeable
- EnPissant 11mo agoCompaction is just what Claude Code has done forever, right?
- enraged_camel 11mo agoI am also trying to understand the difference between compaction, and what IDEs like Cursor do when they "summarize" context over long-running conversations. Is this saying that said summarization now happens at the model level? Or are there other differences?
- GardenLetter27 11mo ago
- causal 11mo agoSigh. Time to try it again I guess. I give OpenAI way more chances than it deserves.
- EcommerceFlow 11mo agoGemini 3 had a great 24 hour SOTA run for coding
- CuriouslyC 11mo agoGemini is still the best oracle/planner by a mile. It's just a bad agent. Give it a bundle of your repo and get it to plan your changes, then hand it off to codex to implement.
- ygouzerh 11mo agoGood idea! I found Gemini have horribly slow for anything
- croes 11mo agoThe new detergent now washes even whiter
- bgwalter 11mo agoCome on folks, this is funny. They also have industrial strength laundromats to go with the detergent.
- pton_xd 11mo agoI love how programming discussions du jour have basically devolved into "really? my socks definitely smell better after using 2 scoops of last month's soap. what spin cycle are you using?"
- SunshineTheCat 11mo agoMy observation has been that Codex tends to hit logical/data-driven/back-end tasks out of the park while doing weird, random nonsense with even simple UI tasks. This could me needing to improve how I phrase my prompts, but it will be interesting to see if it's improved in that arena at all.
- cube2222 11mo agoSomewhat related, after seeing the praise for codex in the Sonnet 4.5 release thread I gave it a go, and I must say, that CLI is much worse than Claude Code (even if the model is great, I’m not sure where the issue really lies between the two). It was extremely slow (like, multiple times slower than Sonnet with Claude Code, though that’s partially on me for using thinking-high I guess) to finish the task, with the back-and-forths being on the order of tens of minutes. Moreover, the context management seems to be really weird. I’m not sure how exactly it works, but - 1. It uses very little tokens / fills up the context slowly (good I guess) 2. Doesn’t seem to actually internalize the contents of files you mention to it, or it edits. #2 here being the main one - I usually context-dump reference code for Claude Code, and it does a perfect job of adhering to codebase patterns and its architecture, while codex was completely ignorant of the existing code style. Moreover, it wrote extremely defensive code, even for code where it wrote both ends itself. All in all, I was really let down after seeing all the praise.
- agentifysh 11mo agosure claude code has better ux but honestly its hard to get any good amount of usage out of the subscriptions vs what codex offers at the same price with claude im constantly hitting rate limits with codex getting substantially more and "slow" isn't really a problem for me as long as it keep working the only complaint i have is that codex itself has usage limited now (Either due to outstanding git issues around tools or by throttling on their end) compared to a few months ago the true magical moment was codex pro letting me run swarms of agents day in day out without any worries about rate limits it truly felt unlimited if claude manages to release a smaller model or some way to deal with the rapidly depleting usage limits (this is the top complaint on reddit and they eventually just stopped allowing threads about it) it would definitely be used more but for now codex is clearly the workhorse and claude used side by side.
- cube2222 11mo agoWell as I said, codex didn’t adhere to codebase standards for me and the code quality was worse (very defensive), so even after waiting longer, results weren’t there for me. But the subscription thing is a non-issue for me as I use the API, and mostly use Claude Code synchronously, with the occasional rare background agent.
- tosh 11mo agoCodex CLI 0.59 got released (but has no changelog text) https://github.com/openai/codex/releases/tag/rust-v0.59.0 https://github.com/openai/codex/releases/tag/rust-v0.59.0
- bgwalter 11mo agoSo they all release before the Nvidia numbers tonight. The real question is: How well can Nvidia hide the circular deals in the books?
- amluto 11mo agoI would love to see all the big players put 1% of the effort they put into model training into making the basic process of paying and signing in suck less. Claude: they barely have a signin system at all. Multiple account support doesn’t exist. The minimum seat count for business is nonsense. The data retention policies are weak. OpenAI: Make ZDR a thing you can use or buy without talking to sales, already. And for those using containers or a remote system or really anything other than local development with the codex CLI, you really really need to fix this bug. I bet Codex could do at least the client part for you! https://github.com/openai/codex/issues/2798 https://github.com/openai/codex/issues/2798 (Hint: Claude Code gets this right by default, despite the fact that everything else about Claude sign-in is a joke.) Google: get all your B2B AI product managers in one room and tell them that they need to make one single product menu on one single webpage with all the pricing on that page and that the Google Cloud people are not permitted to make anything that isn’t actually logically Google Cloud depend on Google Cloud Billing. Your product cannot compete with OpenAI or Anthropic if people need to ask an LLM to figure out what your product is and if your own fancy LLMs can’t give a straight answer. My company pays for a non-Google product primarily because it’s too complicated to pay for the Google product! Right now, trying to use Google’s AI is like trying to ride Bay Area public transit before the Clipper Card.
- atonse 11mo agoAgree 1,000%. I just won’t even waste my time with the google stuff cuz I can’t figure out how to pay with it. And that’s a problem everywhere at google. Our google play account is suspended cuz I can’t verify the company. It won’t let me cuz it says I’m not the owner. I’ve always been the owner of my company. For 18 years. There is no one else. Once some error said make sure the owner email matches your profile in google payments and I was like, what is google payments and where do I even begin with that? I’ve never paid for google play so what does payments have to do with anything? It’s totally random stuff. Get your shit together, google. Make your products and payment systems coherent, rather than it obviously looking like it was designed by a fiefdom full of territorial managers.
- nico 11mo ago
- kytazo 11mo ago500 Internal Server Error.
- morog 11mo agoditto. Also OpenAI vector stores are down right now across the board
- nakamoto_damacy 11mo agoIt’s good but Gemini 3 beats it.
- syntaxing 11mo agoI rarely used Codex compared to Claude because it was extremely slow in GitHub copilot . Like maybe 2-5X slower than Claude Sonnet. I really wish they just made their models faster than “better”
- nartho 11mo agoHave you tried Mistral ? Definitely one of the fastest models
- syntaxing 11mo agoMy employer doesn’t offer/allow anything besides the “traditional” offerings on GitHub copilot.
- theshrike79 11mo agoI've tried Mistral for coding and it seems to be laughably bad every time. Dunno what I'm doing wrong.
- levocardia 11mo agoVery interesting to see the range of peoples' preferences. I would almost always prefer smart over fast; I have all my LLMs to be all-thinking-all-the-time.
- syntaxing 11mo agoIt’s a balance, I haven’t felt like codex provided anything that Sonnet 4.5 didn’t. Why wait longer for getting the same results. Though that does bring up an interesting point. Anecdotally, Sonnet does a lot more grep-ing while Codex reads files straight up. Might be the difference in speed and maybe smarter models will do better. Once this model is on copilot, I can test it out.
- mrguyorama 11mo agoGPT-5 was recently updated to make it more "thinking" and "warmer" or whatever and now a task (semantically compare these two short files) that used to take 5 seconds and reliably produce useful and consistent output now takes 90 seconds to "think" (while it's thinking output makes it pretty clear there is zero thinking happening) and produces a completely differently structured output every single time, making the tool not only slower and more expensive to use, but worse at a simple task that LLMs should be very good at. There's an option to "get a quick answer" and I hoped clicking that would revert to previous performance and instead what it does is ignore that I uploaded two files and asks me to upload the files Literally the only real good task I've found for these dumb things and they still found a way to fuck it up because they need to keep the weirdos and whales addicted. It's now almost easier to go back to comparing these files by eye, or just bite the bullet and finally write a few lines of python to actually do it right and reliably.
- andai 11mo agoSizeable if veracious!
- the__alchemist 11mo agoThis is a tangent: Has anyone noticed that GPT-5.0 at some point started producing much faster, crappier answers, then 5.1 made it slower + better again? (Both in Thinking mode)
- wincy 11mo agoI did notice that, I thought maybe I’d exceeded my thinking requests
- dgfl 11mo agoAbsolutely. Even in extended thinking mode it was thinking for only a few seconds in prompts that used to take minutes. Much faster token/s in any mode and significantly worse, exactly as you describe. It seems like they might still be heavily nerfing / quantizing the models in production a couple weeks before a new release, like they have always (unofficially) done.
- ygouzerh 11mo agoGPT-5 was horrible. It produced AI slop have immense speed, which is quite tough when other coworkers ask to review their PR...
- Narciss 11mo agoHere we go again....
- johnfn 11mo agoI've been using a lot of Claude and Codex recently. One huge difference I notice between Codex and Claude code is that, while Claude basically disregards your instructions (CLAUDE.md) entirely, Codex is extremely, painfully, doggedly persistent in following every last character of them - to the point that i've seen it work for 30 minutes to convolute some solution that was only convoluted because of some sentence I threw in the instructions I had completely forgotten about. I imagine Codex as the "literal genie" - it'll give you exactly what you asked for. EXACTLY. If you ask Claude to fix a test that accidentally says assert(1 + 1 === 3), it'll say "this is clearly a typo" and just rewrite the test. Codex will rewrite the entire V8 engine to break arithmetic. Both these tools have their uses, and I don't think one approach is universally better. Because Claude just hacks its way to a solution, it is really fast, so I like using it for iterate web work, where I need to tweak some styles and I need a fast iterative loop. Codex is much worse at that because it takes like 5 minutes to validate everything is correct. Codex is much better for longer, harder tasks that have to be correct -- I can just write some script to verify that what it did work, and let it spin for 30-40 minutes.
- nico 11mo ago> Claude basically disregards your instructions (CLAUDE.md) entirely A friend of mine tells Claude to always address him as “Mr Tinkleberry”, he says he can tell when Claude is not paying attention to the instructions on CLAUDE.md when Claude stops calling him “Mr Tinkleberry” consistently
- benzible 11mo agoYep, it's David Lee Roth's brown M&M trick https://www.smithsonianmag.com/arts-culture/why-did-van-halen-demand-concert-venues-remove-brown-mms-from-the-menu-180982570/ https://www.smithsonianmag.com/arts-culture/why-did-van-hale...
- awad 11mo agoHighly recommend adding some kind of canary like this in all LLM project instructions. I prefer my instructions to say 'always start output with an (uniquely decided by you) emoji' as it's easier to visually scan for one when reading a wall of LLM output, and use a different emoji per project because what's life without a little whim?
- wilg 11mo agoI have been using GPT 5 High Fast in Cursor primarily over Codex, because Codex seems to take way longer and generally annoy me by doing strange CLI stuff, but hopefully I can switch to this new one. I also tried it against Gemini 3 Pro in Cursor and it's hard to tell but at least in some cases I felt like GPT5 was giving better results.
- LZ_Khan 11mo agoWoah, metr results look impressive. Still looking exponential
- tunesmith 11mo agoI've been dealing with Codex CLI for a while and I love it, but I'm wondering if my thinking is just limited. While I'm starting discussions and creating plan docs, I've never been able to ask it to do anything that takes it longer than 25 minutes or so. Usually far less. I'm having trouble imagining what I can ask it to do that would make it take hours - like, wouldn't that require putting together an absolutely massive planning doc that would take hours to put together anyway? I'd rather just move incrementally.
- GenerWork 11mo agoPerhaps they're combining an incredibly complex product that has a lot of interactive features, a big codebase, test creation, and maybe throwing some MCP stuff in there such as creating creating a ticket in Jira if a test fails?
- CuriouslyC 11mo agoEasy way to get an agent to run a long time is just to get it to babysit CI/CD, tell it to iterate on it until it passes. I got Sonnet 4 to run for >6 hours that way.
- aerhardt 11mo agoThe idea of giving it a task that may take six hours and reviewing it also gives me shivers. I'm a very happy Codex customer, but everything turns to disgusting slop if I don't provide: (1) Up-to-date AGENTS.md and an excellent prompt (2) A full file-level API with function signatures, return types and function-level guidance if it's a complex one (3) Multiple rounds of feedback until the result is finely sculpted Overall it's very small units of work - one file or two, tops. I've been letting the above standards go for the last couple of weeks due to crunch and looking at some of the hotspots of slop now lying around has me going all Homelander-face [1] at the sight of them. Those hotspots are a few hundred lines in the worst cases; I'm definitely not ready to deal with the fallout of any unit of work that takes even more than 20min. [1] https://i.kym-cdn.com/entries/icons/original/000/050/702/ab7.jpg https://i.kym-cdn.com/entries/icons/original/000/050/702/ab7...
- 11mo ago
- 999900000999 11mo agoI really would prefer them to start creating customized models. I've vibe coded Godot games extensively. Just about every model I've tried likes to invent imaginary functions. I was really prefer for there to be a way for me to pick model trained in whatever framework I need. Reviewing AI generated code feels like editing a long book, and every now and then you notice some words are just completely made up. You then ask the AI to fix its book, and it will just add more AI generated words. On one hand I want this to be a reality check to everyone who's trying to lay off real software engineers to replace us with AI. On the other hand half of the stock market is held up by overhyped AI valuations. If the tide goes out too fast, and there is a mass realization that this stuff just isn't as good as it's hyped to be, it's not going to be fun for anyone.
- andai 11mo agoI had this problem 2 years ago. All the models were telling me use libraries that hadn't been invented yet. That was annoying back then, but these days that's not so much of a problem. You can write your program and then simply have it invent the library as well, while it's at it! ;)
- razodactyl 11mo agoThese days not so much of a problem because the libraries now exist? Haha
- karmajunkie 11mo agomostly because of slop-squatting i’d imagine…
- int_19h 11mo agoIt's still very much a problem. For one hilarious example, Gemini (2.5; I haven't tried it with 3 yet) only knows about the old Google API for Gemini, not about the new one. So if you give it code written against the new stuff, it will often do things like, "this is definitely wrong, I know this API doesn't have this method, let me fix that".
- spectraldrift 11mo agoWeird how they only share three hand-picked evals, ignoring the evals where they were left in the dust like ARC-AGI2. This post is so misleading, I don't even know whether to trust the numbers they did share. One is just fraction of a percentage point away from Gemini 3 pro, which is awfully convenient for marketing and easy to hide. Very open, OpenAI.
- XenophileJKO 11mo agoNot really that weird. This isn't intended to be a "general" model. This is a coding model so they showed the coding evals. The assumption would be relative to GPT5.1, non-coding evals would be likely regress or be similar. Like when advertising the new airliner, most people don't care about how fast it taxis.
- simonw 11mo agoThinking level medium: https://tools.simonwillison.net/svg-render#%3Csvg%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%20width%3D%22640%22%20height%3D%22480%22%20viewBox%3D%220%200%20640%20480%22%20role%3D%22img%22%20aria-label%3D%22Pelican%20riding%20a%20bicycle%22%3E%0A%20%20%3Cdefs%3E%0A%20%20%20%20%3ClinearGradient%20id%3D%22sky%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%220%25%22%20y2%3D%22100%25%22%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%23bfe9ff%22%2F%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%23f7fcff%22%2F%3E%0A%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%3ClinearGradient%20id%3D%22sand%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%22100%25%22%20y2%3D%220%25%22%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%23f7d9a6%22%2F%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%23f2c689%22%2F%3E%0A%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%3ClinearGradient%20id%3D%22sea%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%22100%25%22%20y2%3D%220%25%22%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%2371b7e6%22%2F%3E%0A%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%234fa0d6%22%2F%3E%0A%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%3Cfilter%20id%3D%22shadow%22%20x%3D%22-20%25%22%20y%3D%22-20%25%22%20width%3D%22140%25%22%20height%3D%22160%25%22%3E%0A%20%20%20%20%20%20%3CfeDropShadow%20dx%3D%220%22%20dy%3D%226%22%20stdDeviation%3D%226%22%20flood-color%3D%22%231c2b38%22%20flood-opacity%3D%220.25%22%2F%3E%0A%20%20%20%20%3C%2Ffilter%3E%0A%20%20%20%20%3Cstyle%3E%0A%20%20%20%20%20%20.line%20%7B%20stroke-linecap%3A%20round%3B%20stroke-linejoin%3A%20round%3B%20fill%3A%20none%3B%20%7D%0A%20%20%20%20%20%20.wheel%20%7B%20stroke%3A%20%231f2d3a%3B%20stroke-width%3A%208%3B%20%7D%0A%20%20%20%20%20%20.frame%20%7B%20stroke%3A%20%23f45d4c%3B%20stroke-width%3A%208%3B%20%7D%0A%20%20%20%20%20%20.accent%20%7B%20stroke%3A%20%23f3a712%3B%20stroke-width%3A%206%3B%20%7D%0A%20%20%20%20%20%20.feather%20%7B%20fill%3A%20%23f5f7fb%3B%20stroke%3A%20%23d6dce5%3B%20stroke-width%3A%203%3B%20%7D%0A%20%20%20%20%20%20.beak%20%7B%20fill%3A%20%23f9b24c%3B%20stroke%3A%20%23d88a1f%3B%20stroke-width%3A%203%3B%20%7D%0A%20%20%20%20%20%20.eye%20%7B%20fill%3A%20%23ffffff%3B%20stroke%3A%20%231f2d3a%3B%20stroke-width%3A%202%3B%20%7D%0A%20%20%20%20%20%20.leg%20%7B%20stroke%3A%20%23d88a1f%3B%20stroke-width%3A%207%3B%20%7D%0A%20%20%20%20%3C%2Fstyle%3E%0A%20%20%3C%2Fdefs%3E%0A%20%20%3Crect%20width%3D%22640%22%20height%3D%22320%22%20fill%3D%22url(%23sky)%22%2F%3E%0A%20%20%3Crect%20y%3D%22320%22%20width%3D%22640%22%20height%3D%2260%22%20fill%3D%22url(%23sea)%22%2F%3E%0A%20%20%3Crect%20y%3D%22380%22%20width%3D%22640%22%20height%3D%22100%22%20fill%3D%22url(%23sand)%22%2F%3E%0A%20%20%3Cg%20filter%3D%22url(%23shadow)%22%3E%0A%20%20%20%20%3Cline%20class%3D%22wheel%20line%22%20x1%3D%22200%22%20y1%3D%22360%22%20x2%3D%22440%22%20y2%3D%22360%22%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%224%22%20opacity%3D%220.2%22%2F%3E%0A%20%20%20%20%3Cg%20class%3D%22wheel%22%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%22200%22%20cy%3D%22340%22%20r%3D%2260%22%20fill%3D%22%23eef3f9%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%22440%22%20cy%3D%22340%22%20r%3D%2260%22%20fill%3D%22%23eef3f9%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%22200%22%20cy%3D%22340%22%20r%3D%2212%22%20fill%3D%22%23f45d4c%22%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%22440%22%20cy%3D%22340%22%20r%3D%2212%22%20fill%3D%22%23f45d4c%22%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%224%22%2F%3E%0A%20%20%20%20%20%20%3Cg%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%224%22%20opacity%3D%220.7%22%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22280%22%20x2%3D%22200%22%20y2%3D%22400%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22140%22%20y1%3D%22340%22%20x2%3D%22260%22%20y2%3D%22340%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22160%22%20y1%3D%22300%22%20x2%3D%22240%22%20y2%3D%22380%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22160%22%20y1%3D%22380%22%20x2%3D%22240%22%20y2%3D%22300%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22440%22%20y1%3D%22280%22%20x2%3D%22440%22%20y2%3D%22400%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22380%22%20y1%3D%22340%22%20x2%3D%22500%22%20y2%3D%22340%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22400%22%20y1%3D%22300%22%20x2%3D%22480%22%20y2%3D%22380%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cline%20x1%3D%22400%22%20y1%3D%22380%22%20x2%3D%22480%22%20y2%3D%22300%22%2F%3E%0A%20%20%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3Cg%20class%3D%22frame%22%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22340%22%20x2%3D%22300%22%20y2%3D%22280%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22300%22%20y1%3D%22280%22%20x2%3D%22440%22%20y2%3D%22340%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22200%22%20y1%3D%22340%22%20x2%3D%22320%22%20y2%3D%22340%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22320%22%20y1%3D%22340%22%20x2%3D%22340%22%20y2%3D%22300%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22320%22%20y1%3D%22340%22%20x2%3D%22320%22%20y2%3D%22280%22%20stroke%3D%22%23f3a712%22%20stroke-width%3D%226%22%2F%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3Cline%20class%3D%22accent%22%20x1%3D%22440%22%20y1%3D%22340%22%20x2%3D%22470%22%20y2%3D%22280%22%2F%3E%0A%20%20%20%20%3Cline%20class%3D%22accent%22%20x1%3D%22320%22%20y1%3D%22340%22%20x2%3D%22260%22%20y2%3D%22300%22%2F%3E%0A%20%20%20%20%3Crect%20x%3D%22310%22%20y%3D%22260%22%20width%3D%2220%22%20height%3D%2212%22%20rx%3D%224%22%20fill%3D%22%231f2d3a%22%2F%3E%0A%20%20%20%20%3Cg%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%226%22%20stroke-linecap%3D%22round%22%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22320%22%20y1%3D%22340%22%20x2%3D%22290%22%20y2%3D%22370%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22320%22%20y1%3D%22340%22%20x2%3D%22350%22%20y2%3D%22370%22%2F%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3Cg%20stroke%3D%22%231f2d3a%22%20stroke-width%3D%226%22%20stroke-linecap%3D%22round%22%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22470%22%20y1%3D%22280%22%20x2%3D%22500%22%20y2%3D%22250%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22470%22%20y1%3D%22280%22%20x2%3D%22500%22%20y2%3D%22300%22%2F%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%3C%2Fg%3E%0A%20%20%3Cg%20transform%3D%22translate(260%20120)%22%20filter%3D%22url(%23shadow)%22%3E%0A%20%20%20%20%3Cellipse%20class%3D%22feather%22%20cx%3D%2250%22%20cy%3D%22120%22%20rx%3D%2260%22%20ry%3D%2280%22%2F%3E%0A%20%20%20%20%3Cellipse%20class%3D%22feather%22%20cx%3D%22110%22%20cy%3D%2290%22%20rx%3D%2270%22%20ry%3D%2260%22%20transform%3D%22rotate(-10%20110%2090)%22%2F%3E%0A%20%20%20%20%3Cellipse%20class%3D%22feather%22%20cx%3D%2230%22%20cy%3D%2290%22%20rx%3D%2240%22%20ry%3D%2250%22%20transform%3D%22rotate(30%2030%2090)%22%2F%3E%0A%20%20%20%20%3Cpath%20class%3D%22beak%22%20d%3D%22M170%20110%20Q230%20120%20250%20100%20Q230%2090%20180%2075%20Z%22%2F%3E%0A%20%20%20%20%3Cpath%20class%3D%22feather%22%20d%3D%22M120%2070%20Q150%2040%20190%2050%20Q170%2070%20160%2092%20Z%22%2F%3E%0A%20%20%20%20%3Ccircle%20class%3D%22eye%22%20cx%3D%22160%22%20cy%3D%2278%22%20r%3D%2212%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22160%22%20cy%3D%2278%22%20r%3D%225%22%20fill%3D%22%231f2d3a%22%2F%3E%0A%20%20%20%20%3Cpath%20class%3D%22feather%22%20d%3D%22M40%20150%20Q20%20200%2030%20230%20Q60%20200%2070%20170%20Z%22%2F%3E%0A%20%20%20%20%3Cpath%20class%3D%22feather%22%20d%3D%22M80%20150%20Q70%20200%2090%20230%20Q120%20200%20120%20170%20Z%22%2F%3E%0A%20%20%20%20%3Cline%20class%3D%22leg%22%20x1%3D%2240%22%20y1%3D%22200%22%20x2%3D%2230%22%20y2%3D%22250%22%2F%3E%0A%20%20%20%20%3Cline%20class%3D%22leg%22%20x1%3D%2290%22%20y1%3D%22200%22%20x2%3D%22100%22%20y2%3D%22250%22%2F%3E%0A%20%20%20%20%3Cg%20stroke%3D%22%23d88a1f%22%20stroke-width%3D%227%22%20stroke-linecap%3D%22round%22%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%2230%22%20y1%3D%22250%22%20x2%3D%22-5%22%20y2%3D%22260%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22100%22%20y1%3D%22250%22%20x2%3D%22130%22%20y2%3D%22260%22%2F%3E%0A%20%20%20%20%3C%2Fg%3E%0A%20%20%20%20%3Cpath%20class%3D%22feather%22%20d%3D%22M60%2040%20Q30%2010%2020%20-20%20Q80%2010%20110%2040%20Z%22%2F%3E%0A%20%20%3C%2Fg%3E%0A%3C%2Fsvg%3E%0A https://tools.simonwillison.net/svg-render#%3Csvg%20xmlns%3D... Thinking level xhigh: https://tools.simonwillison.net/svg-render#%20%20%3Csvg%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%20viewBox%3D%220%200%20260%20220%22%20aria-label%3D%22Pelican%20riding%20a%20bicycle%22%3E%0A%20%20%20%20%3Cdefs%3E%0A%20%20%20%20%20%20%3ClinearGradient%20id%3D%22sky%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%220%25%22%20y2%3D%22100%25%22%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%23b7d9ff%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%23eef7ff%22%2F%3E%0A%20%20%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%20%20%3ClinearGradient%20id%3D%22beak%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%22100%25%22%20y2%3D%220%25%22%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%23f9b233%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%23f48c00%22%2F%3E%0A%20%20%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%20%20%3ClinearGradient%20id%3D%22frame%22%20x1%3D%220%25%22%20y1%3D%220%25%22%20x2%3D%22100%25%22%20y2%3D%22100%25%22%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%220%25%22%20stop-color%3D%22%231f6fa3%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%230f3e63%22%2F%3E%0A%20%20%20%20%20%20%3C%2FlinearGradient%3E%0A%20%20%20%20%20%20%3CradialGradient%20id%3D%22tire%22%20cx%3D%2250%25%22%20cy%3D%2250%25%22%20r%3D%2250%25%22%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%2260%25%22%20stop-color%3D%22%231b1b1b%22%2F%3E%0A%20%20%20%20%20%20%20%20%3Cstop%20offset%3D%22100%25%22%20stop-color%3D%22%230e0e0e%22%2F%3E%0A%20%20%20%20%20%20%3C%2FradialGradient%3E%0A%20%20%20%20%20%20%3Cstyle%3E%0A%20%20%20%20%20%20%20%20.feather%20%7B%20fill%3A%20%23f7f9fb%3B%20stroke%3A%20%23cdd6e0%3B%20stroke-width%3A%201.6%3B%20%7D%0A%20%20%20%20%20%20%20%20.outline%20%7B%20stroke%3A%20%23123147%3B%20stroke-width%3A%203%3B%20stroke-linecap%3A%20round%3B%20stroke-linejoin%3A%20round%3B%20fill%3A%20none%3B%20%7D%0A%20%20%20%20%20%20%3C%2Fstyle%3E%0A%20%20%20%20%3C%2Fdefs%3E%0A%0A%20%20%20%20%3Crect%20width%3D%22260%22%20height%3D%22220%22%20fill%3D%22url(%23sky)%22%2F%3E%0A%20%20%20%20%3Cellipse%20cx%3D%22130%22%20cy%3D%22190%22%20rx%3D%22110%22%20ry%3D%2214%22%20fill%3D%22%23d9e7f5%22%20opacity%3D%220.7%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Bicycle%20--%3E%0A%20%20%20%20%3Ccircle%20cx%3D%2280%22%20cy%3D%22168%22%20r%3D%2236%22%20fill%3D%22url(%23tire)%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%2280%22%20cy%3D%22168%22%20r%3D%2222%22%20fill%3D%22%233a4b5c%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%2280%22%20cy%3D%22168%22%20r%3D%227%22%20fill%3D%22%23dfe7ef%22%2F%3E%0A%0A%20%20%20%20%3Ccircle%20cx%3D%22182%22%20cy%3D%22168%22%20r%3D%2236%22%20fill%3D%22url(%23tire)%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22182%22%20cy%3D%22168%22%20r%3D%2222%22%20fill%3D%22%233a4b5c%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22182%22%20cy%3D%22168%22%20r%3D%227%22%20fill%3D%22%23dfe7ef%22%2F%3E%0A%0A%20%20%20%20%3Cpath%20d%3D%22M80%20168%20L118%20140%20L158%20140%20L182%20168%22%20fill%3D%22none%22%20stroke%3D%22url(%23frame)%22%20stroke-width%3D%227%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22118%22%20y1%3D%22140%22%20x2%3D%22140%22%20y2%3D%22110%22%20class%3D%22outline%22%20stroke%3D%22url(%23frame)%22%20stroke-width%3D%228%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22140%22%20y1%3D%22110%22%20x2%3D%22182%22%20y2%3D%22168%22%20class%3D%22outline%22%20stroke%3D%22url(%23frame)%22%20stroke-width%3D%228%22%2F%3E%0A%20%20%20%20%3Cline%20x1%3D%22118%22%20y1%3D%22140%22%20x2%3D%22102%22%20y2%3D%22168%22%20class%3D%22outline%22%20stroke%3D%22url(%23frame)%22%20stroke-width%3D%226%22%2F%3E%0A%20%20%20%20%3Crect%20x%3D%22102%22%20y%3D%22130%22%20width%3D%2230%22%20height%3D%226%22%20rx%3D%223%22%20fill%3D%22%230f3e63%22%2F%3E%0A%20%20%20%20%3Crect%20x%3D%22146%22%20y%3D%22102%22%20width%3D%2214%22%20height%3D%2226%22%20rx%3D%226%22%20fill%3D%22%230f3e63%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Pedals%20--%3E%0A%20%20%20%20%3Cg%20transform%3D%22rotate(-30%20118%20140)%22%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22118%22%20y1%3D%22140%22%20x2%3D%22140%22%20y2%3D%22140%22%20stroke%3D%22%23123147%22%20stroke-width%3D%224%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%22140%22%20cy%3D%22140%22%20r%3D%226%22%20fill%3D%22%23f48c00%22%2F%3E%0A%20%20%20%20%20%20%3Cline%20x1%3D%22118%22%20y1%3D%22140%22%20x2%3D%2296%22%20y2%3D%22140%22%20stroke%3D%22%23123147%22%20stroke-width%3D%224%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%20%20%3Ccircle%20cx%3D%2296%22%20cy%3D%22140%22%20r%3D%226%22%20fill%3D%22%23f48c00%22%2F%3E%0A%20%20%20%20%3C%2Fg%3E%0A%0A%20%20%20%20%3C!--%20Pelican%20body%20--%3E%0A%20%20%20%20%3Cellipse%20cx%3D%22130%22%20cy%3D%22115%22%20rx%3D%2254%22%20ry%3D%2236%22%20class%3D%22feather%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M94%20114%20Q80%20118%2080%20134%20Q84%20146%2096%20146%22%20fill%3D%22%23f7f9fb%22%20stroke%3D%22%23cdd6e0%22%20stroke-width%3D%221.6%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M166%20114%20Q180%20118%20180%20134%20Q176%20146%20164%20146%22%20fill%3D%22%23f7f9fb%22%20stroke%3D%22%23cdd6e0%22%20stroke-width%3D%221.6%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M132%2092%20Q140%2078%20160%2070%20Q176%2076%20182%2092%20Q170%2088%20156%2096%20Q142%20104%20132%2092%20Z%22%20fill%3D%22%23f7f9fb%22%20stroke%3D%22%23cdd6e0%22%20stroke-width%3D%221.6%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Wing%20--%3E%0A%20%20%20%20%3Cpath%20d%3D%22M120%20110%20Q94%20118%2088%20142%20Q116%20140%20132%20132%20Q148%20124%20156%20112%20Q134%20114%20120%20110%20Z%22%20fill%3D%22%23e8eef5%22%20stroke%3D%22%23c4ced8%22%20stroke-width%3D%221.5%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Head%20and%20beak%20--%3E%0A%20%20%20%20%3Cellipse%20cx%3D%22168%22%20cy%3D%2288%22%20rx%3D%2226%22%20ry%3D%2220%22%20class%3D%22feather%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M182%2088%20Q214%2086%20224%2096%20Q208%20104%20184%20106%20Q175%20104%20176%2096%20Z%22%20fill%3D%22url(%23beak)%22%20stroke%3D%22%23c77400%22%20stroke-width%3D%221.8%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22172%22%20cy%3D%2282%22%20r%3D%224.4%22%20fill%3D%22%23123147%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22173.5%22%20cy%3D%2280.5%22%20r%3D%221.3%22%20fill%3D%22%23fefefe%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Neck%20--%3E%0A%20%20%20%20%3Cpath%20d%3D%22M150%2094%20Q142%20116%20134%20130%20Q150%20130%20164%20118%20Q170%20110%20168%2098%20Z%22%20fill%3D%22%23f7f9fb%22%20stroke%3D%22%23cdd6e0%22%20stroke-width%3D%221.6%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Legs%20--%3E%0A%20%20%20%20%3Cpath%20d%3D%22M124%20146%20L118%20168%22%20stroke%3D%22%23e1951c%22%20stroke-width%3D%226%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M144%20148%20L152%20168%22%20stroke%3D%22%23e1951c%22%20stroke-width%3D%226%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22118%22%20cy%3D%22168%22%20r%3D%224%22%20fill%3D%22%23d87f00%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22152%22%20cy%3D%22168%22%20r%3D%224%22%20fill%3D%22%23d87f00%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Handlebar%20%2B%20seat%20--%3E%0A%20%20%20%20%3Crect%20x%3D%22104%22%20y%3D%22126%22%20width%3D%2218%22%20height%3D%226%22%20rx%3D%223%22%20fill%3D%22%23123147%22%2F%3E%0A%20%20%20%20%3Cpolyline%20points%3D%22140%2C110%20160%2C92%20178%2C100%22%20fill%3D%22none%22%20stroke%3D%22%23123147%22%20stroke-width%3D%225%22%20stroke-linecap%3D%22round%22%20stroke-linejoin%3D%22round%22%2F%3E%0A%20%20%20%20%3Ccircle%20cx%3D%22178%22%20cy%3D%22100%22%20r%3D%224%22%20fill%3D%22%23123147%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Feather%20details%20--%3E%0A%20%20%20%20%3Cpath%20d%3D%22M110%20118%20Q122%20126%20136%20128%22%20stroke%3D%22%23d6dfe8%22%20stroke-width%3D%221.2%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M114%20126%20Q128%20134%20142%20134%22%20stroke%3D%22%23d6dfe8%22%20stroke-width%3D%221.2%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M118%20134%20Q132%20142%20146%20142%22%20stroke%3D%22%23d6dfe8%22%20stroke-width%3D%221.2%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%2F%3E%0A%0A%20%20%20%20%3C!--%20Breeze%20lines%20--%3E%0A%20%20%20%20%3Cpath%20d%3D%22M32%2070%20Q64%2068%2082%2078%22%20stroke%3D%22%23a4c8e8%22%20stroke-width%3D%223%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%20opacity%3D%220.6%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M24%2092%20Q56%2090%2070%20100%22%20stroke%3D%22%23a4c8e8%22%20stroke-width%3D%223%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%20opacity%3D%220.6%22%2F%3E%0A%20%20%20%20%3Cpath%20d%3D%22M36%20112%20Q60%20112%2078%20120%22%20stroke%3D%22%23a4c8e8%22%20stroke-width%3D%223%22%20fill%3D%22none%22%20stroke-linecap%3D%22round%22%20opacity%3D%220.6%22%2F%3E%0A%20%20%3C%2Fsvg%3E%0A https://tools.simonwillison.net/svg-render#%20%20%3Csvg%20xm...
- andai 11mo agoThe graph showing higher performance for fewer thinking tokens is really interesting! It would be even more interesting to see how Sonnet and Haiku compare with that curve.
- tptacek 11mo agoIs "compaction" a trained-in feature of the model, or just tooling around the model calls? Agents already do compaction.
- kachapopopow 11mo agonot sure if I am actually using 5.1-codex-max or just normal 5.1-codex (is there even 5.1-codex?) trying to continue work where gemini 3 left off and couple prompts in I had to switch back since it was reimplementing and changing things that didn't need changing and attempted to solve typos by making the code implementing those things work with the typo, weird behavior - probably is not compatible with the style gemini tries to solve problems.
- sumedh 11mo agoJust run the /model command in codex and select the model which you want.
- rolisz 11mo agoI got prompted to try it out on the web. It gave me this after 5 minutes: "I wasn’t able to finish creating the new base homepage module template and updating every module to inherit from it within the available time. I did not make any changes or commits." Told it to get back to work. Let's see how that goes.
- epolanski 11mo agoSmall ot question on the GPT cli tool. I gave it a shot last month but I did not enjoy it due to the lack of a proper planning mode and being able to accept each edit independently, has it improved?
- theshrike79 11mo agoNo. Claude is still the only CLI agent tool with a planning mode. Crush, Gemini, Codex and Copilot don't have it for some reason. Can't be that difficult
- hereme888 11mo agoIt's getting so cut-throat for who has the current SOTA model. Seems to be the big income driver.
- kilroy123 11mo agoAll the frontier models seem fairly neck to neck. I wonder which company or lab will finally leapfrog the others with some kind of breakthrough? It sounded like Gemini 3 would be that but in my limit testing it didn't appear to be that.
- boole1854 11mo agoToday I did some comparisons of GPT-5.1-Codex-Max (on high) in the Codex CLI versus Gemini 3 Pro in the Gemini CLI. - As a general observation, Gemini is less easy to work with as a collaborator. If I ask the same question to both models, Codex will answer the question. Gemini will read some intention behind the question, write code to implement the intention, and only then answer the question. In one case, it took me five rounds of repeatedly rewriting my prompt in various ways before I could get it to not code but just answer the question. - Subjectively, it seemed to me that the code that Gemini wrote was more similar to code that I, as a senior-level developer, would have written than what I have been used to from recent iterations of GPT-5.1. The code seemed more readable-by-default and not merely technically correct. I was happy to see this. - Gemini seems to have a tendency to put its "internal dialogue" into comments. For example, "// Here we will do X because of reason Y. Wait, the plan calls for Z instead. Ok, we'll do Z.". Very annoying. I did two concrete head-to-head comparisons where both models had the same code and the same prompt. First, both models were told to take a high-level overview of some new functionality that we needed and were told to create a detailed plan for implementing it. Both models' plans were then reviewed by me and also by both models (in fresh conversations). All three of us agreed that Codex's plan was better. In particular, Codex was better at being more comprehensive and at understanding how to integrate the new functionality more naturally into the existing code. Then (in fresh conversations), both models were told to implement that plan. Afterwards, again, all three of us compared the resulting solutions. And, again, all three of us agreed that Codex's implementation was better. Notably, Gemini (1) hallucinated database column names, (2) ignored parts of the functionality that the plan called for, and (3) did not produce code that was integrated as well with the existing codebase. In its favor, it did produce a better version of a particular finance-related calculation function than Codex did. Overall, Codex was the clear winner today. Hallucinations and ignored requirements are big problems that are very annoying to deal with when they happen. Additionally, Gemini's tendencies to include odd comments and to jump past the discussion phase of projects both make it more frustrating to work with, at this stage.
- jadbox 11mo agoTry checking your temp for any tool using Gemini. "For Gemini 3, we strongly recommend keeping the temperature parameter at its default value of 1.0.While previous models often benefited from tuning temperature to control creativity versus determinism, Gemini 3's reasoning capabilities are optimized for the default setting. Changing the temperature (setting it below 1.0) may lead to unexpected behavior, such as looping or degraded performance, particularly in complex mathematical or reasoning tasks." https://ai.google.dev/gemini-api/docs/gemini-3?thinking=high https://ai.google.dev/gemini-api/docs/gemini-3?thinking=high
- atonse 11mo agoI just tried this out, and was VERY impressed with the speed of the plan mode. I was also totally fine with the code it wrote. Then I made the mistake of saying "run npm run build and fix all issues" (something I've run probably 50 times across codex and cc in the past 2 months). CC does it pretty much 100% of the time. I walked away from Codex, and when I came back, it had installed 2 new node packages, and gone down some crazy rabbit hole with eslint and something else. (this was for 2 minor typescript errors) After I reverted all its changes, had CC do it and it fixed it in about 30-60 seconds. I'll try a few more times. Let's see.
- ansc 11mo agoWhat's the plan mode?
- atonse 11mo agoSorry I mis-worded that. It was my BRAIN being in plan mode (I know CC has a plan mode). I usually ask it to come up with a plan for doing X, and then wait a while for it to look at the code, etc. But in some odd way, GPT-5.1-Codex-Max came up with a plan within 5 seconds. I just found that surprising.
- esafak 11mo agoHow efficient is it; does it go through your subscription quota faster?
- nowittyusername 11mo agoGlad to see evolution of proper context management. the automatic compacting is months overdue so happy to see it finally come.
- ed_mercer 11mo agoAs a long time CC user, I was like "Wait, they didn't have auto-compaction all this time??"
- freediver 11mo agoFirst time that there is a worthy alternative to Claude Code. Codex Max solved a problem I had Claude Code fail multiple times. Gemini CLI was never a contender (between log in/activation/rate limits - wth), will say though that Gemini CLI has the nicest terminal UI.
- deleted 11mo ago[deleted]
- AIorNot 11mo agoAnyone compare this to sonnet 4.5 on full stack development yet
- jwpapi 11mo agoI really hope one day Ill work on challenges that need these new type of agents. Currently, I either need a fast agent that does what I want faster than I can type it (CRUD, forms, etc) or I need an agent to discuss a plan, ups and downs. Whenever I try to give it a bigger task it takes a lot of time, and often is not what I’ve expected, which might be totally my fault or context specific, but as soon as I’m able to define the task properly I would prefer a faster model as it will be good enough, but faster. I really don’t have problems anymore that I can’t reasonable solve fast enough with this approach. I’ve run multiple gpt-5 codex concurrent sessions in the cloud, but I didn’t accept one thing they did. Eventually thinking through it, reading hack boom is faster than outsourcing the work for 30 minutes + 30 minutes to digest +30 minutes to change..
- spruce_tips 11mo ago100% agree. composer-1 really has been the sweet spot for me of capability, reliability, and speed. i dont ask it to do too much at once, and this approach + its speed, materially speeds my work up. i generally find i get the most out of models when i feel like im slightly underutilizing their capabilities. the term i use for this is "staying in the pocket"
- jwpapi 11mo agoIs it available via api? Cant find it on openrouter...
- spruce_tips 10mo agoit's in cursor only
- arresin 11mo agoThat’s the bet cursor took with composer 1. It’s dumb but very fast and that makes it better
- the_duke 11mo agoThe key is learning how to provide proper instructions. Treat it as a developer that just joined the project and isn't aware of the conventions. Provide hints for the desired API design, mention relevant code locations that should be read to gain context on the problem, or that do similar things. An AGENTS.md that explains the project and provides some general guidelines also helps a lot. Codex can be incredibly strong when prompted the right way.
- highfrequency 11mo agoIs GPT-5.1-Codex better or worse than GPT-5.1 (Thinking) for straight up mathematical reasoning (ie if it is optimized for making code edits)? Said another way: what is the set of tasks where you expect GPT 5.1 to be better suited than GPT-5.1 Codex? Is it non-coding problems or non-technical problems?
- NickFORGE 11mo agoWe’ve been experimenting with a similar idea but in a browser-native environment — running real containers + a WebSocket terminal + multi-agent workflows. GPT-5.1 (Codex Max especially) seems to handle multi-step refactors a lot more cleanly, and chaining it through CLI agents has been surprisingly reliable. Curious if anyone else is trying agent orchestration beyond the editor itself?
- LordIsBack 10mo agoI don't think GPT-5.1-Codex-Max is better than Sonnet 4.5 still.
- LordIsBack 10mo agoh