8 ms·
Interesting that compaction is done using an encrypted message that "preserves the model's latent understanding of the original conversation": > Since then, th
by westoncb 8mo ago
Interesting that compaction is done using an encrypted message that "preserves the model's latent understanding of the original conversation":
> Since then, the Responses API has evolved to support a special /responses/compact endpoint (opens in a new window) that performs compaction more efficiently. It returns a list of items (opens in a new window) that can be used in place of the previous input to continue the conversation while freeing up the context window. This list includes a special type=compaction item with an opaque encrypted_content item that preserves the model’s latent understanding of the original conversation. Now, Codex automatically uses this endpoint to compact the conversation when the auto_compact_limit (opens in a new window) is exceeded.
- icelancer 8mo agoTheir compaction endpoint is far and away the best in the industry. Claude's has to be dead last.
- kordlessagain 8mo agoYes, agree completely.
- nubg 8mo agoHelp me understand, how is a compaction endpoint not just a Prompt + json_dump of the message history? I would understand if the prompt was the secret sauce, but you make it sound like there is more to a compaction system than just a clever prompt?
- EnPissant 8mo agoMaybe it's a model fine tuned for compaction?
- FuckButtons 8mo agoThey could be operating in latent space entirely maybe? It seems plausible to me that you can just operate on the embedding of the conversation and treat it as an optimization / compression problem.
- e1g 8mo agoYes, Codex compaction is in the latent space (as confirmed in the article): > the Responses API has evolved to support a special /responses/compact endpoint [...] it returns an opaque encrypted_content item that preserves the model’s latent understanding of the original conversation
- xg15 8mo agoIs this what they mean by "encryption" - as in "no human-readable text"? Or are they actually encrypting the compaction outputs before sending them back to the client? If so, why?
- e1g 8mo ago"encrypted_content" is just a poorly worded variable name that indicates the content of that "item" should be treated as an opaque foreign key. No actual encryption (in the cryptographic sense) is involved.
- xg15 8mo agoAh, that makes more sense. Thanks!
- EnPissant 8mo agoAre you sure? For reasoning, encrypted_content is for sure actually encrypted.
- e1g 8mo agoHmmm, no, I don't know this for sure. In my testing, the /compact endpoint seems to work almost too well for large/complex conversations, and it feels like it cannot contain the entire latent space, so I assumed it keeps pointers inside it (ala previous_response_id). On the other hand, OpenAI says it's stateless and compatible with Zero Data Retention, so maybe it can contain everything.
- Art9681 8mo agoTheir models are specifically trained for their tools. For example the `apply_patch` tool. You would think it's just another file editing tool, but its unique diff format is trained into their models. It also works better than the generic file editing tools implemented in other clients. I can also confirm their compaction is best in class. I've imlemented my own client using their API and gpt-5.2 can work for hours and process millions of input tokens very effectively.
- swalsh 8mo agoIs it possible to use the compactor endpoint independently? I have my own agent loop i maintain for my domain specific use case. We built a compaction system, but I imagine this is better performance.
- jswny 8mo agoHow does this work for other models that aren’t OpenAI models
- westoncb 8mo agoIt wouldn’t work for other models if it’s encoded in a latent representation of their own models.