4 ms·
Pretty basic. The codex app with one conversation per project and several running simultaneously all hours. I’m going for max caching that way and it never gets
by CompoundEyes 2mo ago
Pretty basic. The codex app with one conversation per project and several running simultaneously all hours. I’m going for max caching that way and it never gets lost even with compaction somehow. Each has a plan with milestones to keep up to date and a thin agents file. I check in on them in the Remote app. Use case is protocol and control reverse engineering of audio hardware. I think they must be identifying the heavy use agent sessions and cranking up their cache lives so it’s not a big deal for them.
- rpdillon 2mo agoA billion tokens a day is 11,000 tokens a second sustained. How many tokens per second are you getting off of GPT 5.6 Sol per project?
- sheepscreek 2mo agoUltra mode spins up many sub-agents. On a particularly challenging task, I’ve had as many as 29 agents working at one time. Also if you don’t specify, most end up being the same as the parent model which is pretty wasteful. I engineered a skill that spins up Terra High agents for most sub-agents, resorting to Sol Medium for technical research and Luna High for code/in-project research tasks. On a slightly different topic, Luna Max is incredibly capable and doesn’t use as much quota (Luna tokens are dirt cheap).
- iJohnDoe 2mo agoMost importantly, are you seeing a return on investment for time and ultimate outcome? No one can judge the enjoyment, learning, and hobby aspects. Just wondering if there is an end goal for that much overall expenditure (time, money, energy, etc.)
- _zoltan_ 2mo agoI absolutely see a return. I do not pay for this (we have a corporate gateway), for the price of a junior developer I can get 3-4-5 senior developer's work done. it's insane value.
- jeffmcjunkin 2mo agoOften people are counting all tokens, including cached input tokens, for those more impressive "billions of tokens" quotes.
- rpdillon 2mo agoAh, thanks, I'd missed that nuance in the other reply!
- childintime 2mo agoI'd like to see a benchmark on this specific topic: Reverse engineer the hardware protocol from a driver, or just migrate a driver from one OS to another.
- CompoundEyes 2mo agoI’ve chipped on it with each model since 5.2 but 5.6 sol is something else. When it first came out I’d get some refusals but they’ve since stopped. I wonder what an ideal candidate benchmark task would be for that?
- brynnbee 2mo ago5.6 has been a huge pivotal change in reverse engineering tasks for me too (largely extracting game assets from binary client files). Something I spent literally weeks on in January with claude models at the time was solved in about 30 minutes with 5.6 Sol at medium just yesterday. It's both extremely satisfying but at the same time also a little annoying how much time I had previously spent on it only for it to be solved so quickly now. I suspect improvements in AI will continue this happy but annoyed trend.