4 ms·
This title is easy to misinterpret. If I understand correctly: Codex now encrypts sub-agent prompts and hides those prompts from the user. edit: originally was
by niam 3mo ago
This title is easy to misinterpret. If I understand correctly: Codex now encrypts sub-agent prompts and hides those prompts from the user.
edit: originally was "Codex starts encrypting prompts, uses cyphertext for inference instead"
- bartread 3mo agoI imagine this will be because a decent chunk of the IP in Codex is probably within its prompts, how they're built, and how they're sequenced and orchestrated, rather than in the codebase per se. We had this discussion a few months ago where we talked about allowing people to choose an AI provider and provide their API key, thinking about enterprises with "preferred" (read: mandated) AI suppliers. We also wanted to offer the kind of very simple pricing that this is one way of enabling. But we realised pretty quickly that this would/could lead to leaking our back end prompts to customers and, although those prompts are only a part of the value add, if you could build a detailed trace of them then you'd be able to relatively easily reverse engineer a lot of what we're doing. So we quickly dropped that idea.
- saidnooneever 3mo agothe trick about agentic systems is definitely how to do the prompting. things like automation and sandboxing are trivial in comparisson. if you generally ask via API model directly you can see what basic answers it actually yields and how fine tuning prompts and refinements to output as well as adversarial prompts etc are important to get relatively solid results. a lot of expertise of certain domains' workflows is needed to make it functional within that domain. some of this can be yielded via prompting too etc so its also baoance of how much to prompt it vs. how much of it you wanna let it reason over itself. (if you tell it too much i lock it into a path and if you tell too little it will give incomplete results )
- dmurray 3mo agoPerhaps AI providers should support this natively: the customer supplies the API key but doesn't get access to the transcripts.
- bartread 3mo agoI don't know how you'd enforce that unless it was something you could mandate at the level of the API call, and then the API call is rejected if the customer hasn't configured it for "no transcript". It sort of feels like an area of friction even still.
- dmurray 3mo agoI was thinking of an API key that was scoped both to a specific customer and a specific service provider (perhaps both have to do something to provision it). Billing goes to the customer, debug logs etc go to the service provider.
- agumonkey 3mo agoI'm unable to understand how much value can be in low-definability non deterministic prompts. It feels like the kept the right divinity spell into a chest.
- hnlmorg 3mo agoI don’t disagree with your divinity spell comparison but unfortunately there is a lot of value in the prompts because these spells are the “programming languages” of LLMs.
- agumonkey 3mo agoyeah i get it too, i'm just flabbergasted that this is today's market it reminds me of the pre-vulkan game programming days.. drivers were black boxes, game developpers had to resort to magic tricks to do stuff, until everybody got fed up and wanted some logical ground to operate
- bartread 3mo agoIt's a brave new world, etc., isn't it? One does find oneself slightly askance at one's own thinking sometimes, that's for sure. But I suppose, is it really so different? I mean, back in the day moreso than now, a lot of the valuable IP in any system was in the design and specification of that system - the problems usually solved within the design and specificaion (use X algorithm, etc.) - and the code was "just" the implementation of those solutions. So perhaps it's more of a regression in some ways: the value is in the specification (the prompt) once again. Your point about stochastic behaviour is well made though, and there is no way to 100% guarantee or formally verify the behaviour of a system that relies on an underlying technology whose behaviour is fundamentally stochastic.
- oblio 3mo agoFurther proof that this tech stack is immature and would have needed to bake for a more years. In an ideal world this would have been public tech like ARPANET or WWW and there would have been 2-3 major iterations (until the equivalent of Claude 7-8) and only then would everyone have tried to build huge businesses on top of it. I mean, sure, it's sort of usable, but the churn is insane. And we're burning the planet (and probably the economy, too) for it.
- postalcoder 3mo agoIt's also not the first time Codex started encrypting stuff. Their excellent compaction endpoint has served up a giant encrypted blob since at least five months ago.
- hyperbovine 3mo agoFeels strongly like we're in the late-stage-AI-unicorn phase. If this is really their moat then the Chinese companies will win.
- anon373839 3mo agoThis and the overly stylized model names. Mythos? Sol? Please, it's another version bump.
- oblio 3mo agoAnthropic at least is consistent in their naming. They're all literary genres, ever bigger ones.
- literalAardvark 3mo agoOpenAI's are all unnamed or suns, so it's kinda the same picture.
- oblio 3mo ago> OpenAI's are all unnamed or suns, so it's kinda the same picture. Umm, no: > GPT-1, GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o, GPT-4o mini, o1-preview, o1-mini, o1, o3-mini, o4-mini, o3, o3-pro, GPT-4.1, GPT-4.1 mini, GPT-4.1 nano, GPT-5, GPT-5.1, GPT-5.2, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5.5, GPT-5.5 Pro, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol Besides the constant shifting of nouns and adjectives... Luna is the Moon and Terra is Earth.
- 3mo ago
- imhoguy 3mo agoYeah I thought "wow, some homomorphic encryption* stuff", but then "nah, usual greed". * https://en.wikipedia.org/wiki/Homomorphic_encryption https://en.wikipedia.org/wiki/Homomorphic_encryption
- embedding-shape 3mo agoThe title was fixed like 40 minutes ago, when you come back to old browser tabs you probably want to hit that reload button before leaving a comment ;)
- ofjcihen 3mo agoUnnecessarily snide remark for someone just commenting on their interpretation.
- embedding-shape 3mo agoAt that point 50% of all comments were about the title and it had been updated almost a whole hour before parent made their comment. Sorry for being low on patience.
- hansihe 3mo agoIt seems likely to me this was driven by the `ultra` mode in 5.6, which fans subagents to do work. This mode was previously only available in the web UI (what was previously known as pro?) It seems possible they trained this by doing full RL rollouts of agents interacting with each other. They likely view these prompts somewhat the same as raw reasoning traces, they don't want people to train directly on them. I am unsure if this has been confirmed, but there are some signs that the opaque "compaction blob" they return from their dedicated compaction endpoint might not be text at all, rather a latent space representation of the conversation. The fact that OpenAIs compaction seems to be much higher fidelity than a lot of other providers makes me inclined to believe this. If this is true, it doesn't seem far fetched to infer that they might be applying similar techniques to prompting subagents. I would be curious to see if this way of spawning subagents (encrypted blob) is used when subagents of a different model type is spawned.
- wren6991 3mo agoI think you hit the nail on the head here. Having subagent dispatch in the loop for RLVR is something we've already seen in open models, like Kimi K2.5 and later, so it's no great stretch to assume OpenAI are doing it too. If you keep RL'ing the dispatch then the prompts are likely to keep diverging from the type of prompt a person would write (like CoT becoming increasingly incomprehensible), and that divergence is part of their competitive advantage. > rather a latent space representation of the conversation Student/teacher models derived from the same checkpoint convey a lot of latent information through token choice, as in: https://techxplore.com/news/2026-04-ai-chatbot-student-owls.html https://techxplore.com/news/2026-04-ai-chatbot-student-owls.... I wonder if this is something they can take advantage of by training on compaction inside of the RLVR loop?
- tpurves 3mo ago"Latent space representation" I have been waiting for this moment in the evolution of AI. Well, waiting with some trepidation. It seems inevitable that frontier AI's will, at some point, leave behind human-comprehensible representations of language. Purely for functional reasons, it's going to start making sense for AI agents to communicate amongst themselves in much more efficient ways than borrowing the languages of flesh-bag humans as an interface medium. I Imagine next that programming languages, interfaces and API design starts going this direction next. Being written, expressed and optimized as blobs of high dimensional vector space. As humans we might still be able to understand some abstractions of what our AI's are talking about to each other, but maybe not more so then we understand how different regions of our own brain communicate with each other.
- themgt 3mo agoIt's sort of insane though, you not only have dozens/hundreds of stochastic agents running on your machine, but you cannot even inspect the instructions those agents are working off of? I've gone in to look at Claude subagent/workflows and sometimes been like "no this was a mistake to spin up" ... Codex users just get to token yolo the encrypted telephone operator instructions+shell from orchestrator to subagents?
- Jean-Papoulos 3mo agoYou already have an agent freely doing stuff on your machine. Subagents prompts are a weird place to draw a line. It's not like you're reading everything the agent is doing in any case, let's not kid ourselves.
- embedding-shape 3mo agoWhen things go wrong I very much read the session traces to figure out what in my prompt wasn't good/explicit enough, then retry to evaluate if it would have helped. I was about to do the same with Sol + Ultra, but then discovered this encryption issue that prevents me from doing the same for sub-agents.
- realusername 3mo ago> It's not like you're reading everything the agent is doing in any case Personally I do, these tools aren't mature enough to be used without supervision
- halfcat 3mo ago> You already have an agent freely doing stuff on your machine No. Agents run in VMs. Assume anywhere you’re running an agent will be compromised, because eventually, it will be. The only reason most people haven’t is luck, they didn’t happen to install Axios or Tanstack at a certain time.
- djeastm 3mo ago>but you cannot even inspect the instructions those agents are working off of? It makes more sense when you realize they don't want developers to be doing any coding at all. That's what they seem to be moving towards. From product manager to product via AI.