4 ms·
This model is great at long horizon tasks, and Codex now has heartbeats, so it can keep checking on things. Give it your hardest problem that would take hours w
by vthallam 5mo ago
This model is great at long horizon tasks, and Codex now has heartbeats, so it can keep checking on things. Give it your hardest problem that would take hours with verifiable constraints, you will see how good this is:)
*I work at OAI.
- dandaka 5mo agoCould be a great feature, can't wait to test! Tired of other models (looking at you Opus) constantly stuck mid-task lately.
- winrid 5mo agoInteresting, I just had opus convert a 35k loc java game to c++ overnight (root agent that orchestrated and delegated to sub agents) and woke up and it's done and works. What plan are you on? I'm starting to wonder if they're dynamically adjusting reasoning based on plan or something.
- gck1 5mo agoI'm on max 5x and noticed this too. I don't use built-in subagents but rather full Claude session that orchestrates other full claude sessions. Worker agents that receive tasks now stop midway, they ask for permission to continue. My "heartbeat" is basically "status. One line" message sent to the orchestrator. Opus 4.6 worker agents never asked for permission to continue, and when heartbeat was sent to orchestrator, it just knew what to do (checked on subagents etc). Now it just says that it waits for me to confirm something.
- winrid 5mo agoWeird. I don't have this behavior, although I did with codex and 5.4 haha. I bet the providers are playing with settings underneath and different users are routed to different deployments, or they're secretly routing us to different models under load.
- adamandsteve 5mo agoThis has to be bait.
- azan_ 5mo agoWhy?
- winrid 5mo agowhat?
- adamandsteve2 5mo agoBecause there’s no way in hell it can rewrite a game with 35k loc perfectly lol, link the codebase or it didn’t happen.
- winrid 5mo agoI'll be launching the game in a couple months so you can see it then :) There are bugs, it doesn't work perfectly, but that's just part of testing and refinement at this point. My initial prompt was just: "let's work on converting this java game to c++ using panda3d. you're a panda3d c++ expert. you will be the agent that owns the project, creating the plan, and the delegating each step to sub-agents that create each system in the correct order." it created like 17 different tasks and sub agents and opus 4.7 orchestrated it. I did personally validate which rendering engine would be good for the project etc first.
- frotaur 5mo agoI've been using the /ralph-loop plugin for claude code, works well to keep the model hammering at the task.
- dannyw 5mo agoIt's genuinely so great at long horizon tasks! GPT-5.5 solved many long-horizon frontier challenges, for the first time for an AI model we've tested, in our internal evals at Canva :) Congrats on the launch!
- brcmthrowaway 5mo agoCan we not do growth hacking here?
- smallerize 5mo agoHN is owned by a startup accelerator and venture capital firm. They do growth hacking on the front page. And you probably know that since your throwaway account is several years old.
- deleted 5mo ago[deleted]
- RALaBarge 5mo agoWe totally agree. That's what I've been heads down, HUNGRY, working on, looking for investors and founding engineers pst: https://heymanniceidea.com https://heymanniceidea.com (disclaimer: I am not associated with heymanniceidea.com)
- bkyan 5mo agoSorry, what is "heartbeats", exactly?
- thereeldeel 5mo agoWill Codex App support new context window, rather than compaction, for "unrelated" sub-tasks during long horizon tasks?
- rasputin243 5mo agoUse subagents
- spaceman_2020 5mo agoIs there any task that actually doesn't require human intervention in-between, even if its just to setup stuff? Like I will get Opus to make me an app but it will stop in between because I need to setup the db and plug in the API keys and Opus really can't do that on its own yet
- stingraycharles 5mo ago> Is there any task that actually doesn't require human intervention in-between, even if its just to setup stuff? The goal is none. The current situation: everything that matters requires human intervention. I think the end situation will be that LLMs will be able to perform decently well in a highly controlled and predictable environment.
- leodavi 5mo ago> in a highly controlled and predictable environment Why this constraint? A common sentiment I see online (sorry, to group you in) is "[tool] will be capable, actually, but only in a context that trivializes its usefulness." I think modern post-training like RLVR + inference-time output token scaling can _probably_ scale so the agents can solve any computable task, even when placed in noisy or misconfigured environments. But it won't be economical for a long while. But it already seems largely capable of that today.
- furyofantares 5mo agoPorting large amount of code.