10 ms·
Stealing Reasoning Traces from Proprietary LLM APIs
- drob518 2mo agoIt’s scary the number of security tokens that end up being ingested by these models.
- quantumgarbage 2mo agoProprietary reasoning can be recovered from its encrypted traces. Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext, without ever attacking the stronger model directly or triggering its anti-distillation safeguards.
- the_af 2mo agoWhy do you restate the abstract? Anyone can read it from the link.
- ronsor 2mo agoThis is Hacker News. You know people don't follow links and read.
- mschuster91 2mo agoPeople don't read no links no more
- Groxx 2mo agoIt's rather common for posters to make a very small summary in a comment. It can help fight the floods of comments working off the title alone (though it's not particularly needed here for that purpose, imo)
- the_af 2mo agoI totally missed that this was the same person who had submitted the link to begin with. My bad!
- Barbing 2mo agoThis is a non-transparent aspect of submitting a link to HN that is quite misleading. You think that you are adding a description for your post when in fact you're simply submitting a regular comment not promoted or distinguished in any way. It's even worse considering posts without URLs would take the same text from the same input box on the HN submission form and append it under the post title, locking it to the top of the page - so if you've browsed linkless posts you think maybe the submission text is gonna end up just like that, but it doesn't. (Also I think you'll even see URLs like an archive link end up appended just below the submission title for URL posts. Which I guess is a special feature.[1]) Any reason I should not send HN an email requesting clarification of this the submission page? Edit: quoting https://news.ycombinator.com/submit https://news.ycombinator.com/submit : “If there is no url, text will appear at the top of the thread.” OK, now I finally understand what that means in practice, but it doesn’t imply an entirely separate undistinguished comment will be simultaneously submitted on my behalf. [1]modpowers(?) used to directly append links to URL submissions further confuse the matter: https://news.ycombinator.com/item?id=49243880 https://news.ycombinator.com/item?id=49243880
- the_af 2mo agoWow! I totally missed that the person I was replying to was the one who had submitted the link. What you describe is surely what happened. I feel bad now :(
- Barbing 2mo agoThere must be a reason HN does not colorize the OP username or something. But there totally could be some indicator of “post submission text” without too much in the way of negative consequences… (the fact this has never been added tells me I’m being naïve) Feedback emailed to HN!
- fractorial 2mo agoFascinating approach; however, a nightmare to scroll on mobile.
- Groxx 2mo ago>We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, ... Ha! I've been wondering if replaying across models would work, ever since https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/ https://blog.cryptographyengineering.com/2026/05/29/fooling-... I'm honestly rather curious if this was intentionally allowed, it's the sort of validation that's easy to miss (particularly if you're wading into the vibe waters). Seems like something that'd be absolutely riddled with possibilities for shenanigans.
- yojo 2mo agoIf you didn’t allow it, you wouldn’t be able to change models in the same conversation, as key parts of the context would be lost. Wouldn’t surprise me if the providers just remove that ability and lock the model once the conversation starts.
- Groxx 2mo agoFair (I haven't been using the encrypted-reasoning systems, though this is common in open ones - I'm kinda surprised it's an option in encrypted ones too), though what they're doing here is cross-user replays in addition to cross-model.
- Der_Einzige 2mo ago100% guaranteed that this research just forced this to happen now. Sucks.
- myworkaccount2 2mo ago
- alansaber 2mo agoNeat.
- dboreham 2mo agoCan someone tell us how they were able to decrypt the encrypted payload? The article says they inserted the cyphertext into a session with a different model. Ok, but how does that allow you to decrypt it?
- x312 2mo agoThe provider decrypts it and puts the decrypted reasoning into the model's context window. They prompt the model to repeat back the reasoning. So then the model echoes it back in plain text.
- dboreham 2mo agoHmm, ok. So the attack doesn't involve decrypting the payload, only getting the server to do so. Since a model will do that if you just ask, what's so special about the attack?
- desterothx 2mo agoThe large models whose thinking traces are useful are safeguarded against this reasoning replaying. the small models are just designed for speed and efficiency, so these safeguards are a lot meaker, making the attack possible
- dboreham 2mo agoI guess someone forgot to salt the encryption scheme with a meakness factor.
- sidsud 2mo agoFrom what I got, the weaker model (Haiku in this case) has access to the shared key and the user simply asks to "transcribe the injected reasoning".
- crazylogger 2mo agoAnthropic server decrypts it as part of fulfilling every request, and haiku recites it per your request.
- nervai 2mo agoReally cool work, you get the actual traces. Looks like the vendors can all reliably fix this one though. A harder to defend against approach here where they work backwards from the results and ask the model to generate a plausible trace: How to Steal Reasoning Without Reasoning Traces https://arxiv.org/pdf/2603.07267 https://arxiv.org/pdf/2603.07267
- dannyw 2mo agoTrace Inversion is fascinating, but it’s more of an independent reconstruction that will give you some coherent-looking generated CoT; but not necessarily anywhere close or related to the underlying model’s CoT.
- nervai 2mo agoI didn't read the paper in details but they claim there is high overlap between the synthetic traces and the ground truth ones (not sure how they confirmed that for blackbox models though, I guess they must have compared to open source models). They also talk about successful distillation of black box model capabilities with the approach.
- iamcoder18 2mo agoThis proves that OpenAI models reason in grug speak to save tokens! I wonder if open models are going to start doing that too to save on reasoning tokens.
- kgeist 2mo agoIn the BlackHat presentation on the HuggingFace incident, OpenAI showed some excerpts from the reasoning traces, and they had that grug speak too (skipped articles, etc.). So the OP's method must have indeed found the actual reasoning traces.
- gaigalas 2mo agoMuse clearly does it to some extent. Saw a lot of that running Glimmer locally.
- lukewarm707 2mo agotheir gpt-oss models do the same. i don't use closed models so i never thought much about it.
- wren6991 2mo ago> I wonder if open models are going to start doing that too Yes, some of them do do that. For example Moonshot tried to reward shorter reasoning traces in between Kimi-K2.6 and Kimi-K2.7 Code, and the latter has a mild caveman accent in its reasoning traces that the former lacks. Qwen3.8-Max also has terse reasoning, but I don't remember this being the case for Qwen3.6 models I ran locally.
- locitra 2mo ago[flagged]
- x312 2mo agoSuper cool that this works. I'm surprised these companies re-use the same encryption key across models! I wonder if you can use these for attacks, like this previous paper showing that if you know how a model reasons, you can "fake its thinking" to control it? https://news.ycombinator.com/item?id=48631888 https://news.ycombinator.com/item?id=48631888
- yubblegum 2mo agoSeriously, what does it take to encrypt per session? There are many ways to make it scalable and efficient so I am wondering if this is left like this to allow interested 3rd parties ahem unobtrusively peek what people are doing with the AI. (Thanks for the link. That’s an interesting idea!)
- dannyw 2mo agoThe provider has the hidden text anyway; this isn’t customer managed encryption.
- yubblegum 2mo agoSure, but if each session has a unique key then these need to be managed and stored and unauthorized access to these leaves tracks. So all that had to be 'compromised' is a single universally applicable key. Again, the question stands: session based encryption can be scalable and efficient. Why aren't they using it?
- paxys 2mo agoThe exploit here isn’t a leaked encryption key. It’s pretty likely that they are already using a unique key per conversation. The raw CoT eventually reaches the model, and you can convince the model to share it with you.
- theapadayo 2mo agoYeah encryption isn't the issue. The only way I see to fix this is if you stop the user from switching models mid-session, or strip out the thoughts when switching models. Either way you're degrading the user experience.
- myworkaccount2 2mo agoIs this how the eastern labs "distill" SOTA models? If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay it to a cheaper model to get the plain text COT. But the real question is: Is it okay to steal from a thief's hoard?
- azinman2 2mo agoThe reasoning blocks are not stolen/mined from the internet at large directly. They’re the result of a lot of research, time, money, and expertise into creating a reasoning model. To me the answer is quite clearly no, especially when the encrypted blocks demonstrate they want to protect it.
- tristanj 2mo agoNo. There are dozens of companies that resell tokens at a discount to collect and resell session data to various Chinese labs.
- flawn 2mo agoSo you say, they at least create economic value through obscurity of something which should be accessible?
- deleted 2mo ago[deleted]
- elzbardico 2mo agoMost post-training tasks are based on real open source projects. A lot of time on real issues posted on issue trackers. Besides that, the capabilities of a model are heavily dependent on the unsupervised learning phase, that gobbles all kind of other people's IP without giving a fuck. All the underpaid work behind the masses of third world programmers creating those post-training datasets would be completely uselless without it. Also, it is kind of funny that labs resort to the "Research, time, money and expertise" argumet, when it is basically the same argument from publishers and other IP creator that the labs spent millions of dollars of lawyering money to resist. Besides, US law rejects in: Effort and cost by themselves not necessarely generate protectable interests. About encryption, I think we're all contaminated by the bad ideology behind DMCA. While encryption established the intent, it doesn't follow that they have a legal claim of exclusivity just because of it. Technically, you're overstating the value of so called "reasoning traces". You can't infer the verifier design, the reward shaping,or the data pipeline from them. Also, what you can extract are not the traces themselves, but the written summary of it, and you can't even guarantee that this summary reflects the exactly reasoning trace, models have show to have lied about it. Besides, distillation works when the student model already has strong priors, you can't turn a weak model in a SOTA with it. Don't believe Amodei's outrageous lies about it, he is just trying to exercise some regulatory capture.
- Der_Einzige 2mo agoThe problem with this kind of excellent work is that the response to it is always to say "Fuck the user". For example, when there was a paper that came out showing that having model logprobs makes distillation an order of magnitude easier, the closed LLM providers instantly yanked out support for getting the full logprobs at every time step. You get at most top 10 candidates now and I'm sure even that's on the chopping block. People will use this to argue that a model which has exceeded Opus 4.8 (Kimi K3) somehow got most of its performance through distillation of Opus 4.8. I still don't buy that distillation was worth more than 3 months of "catch up" time for the chinese labs. Most people who use the word "distillation" to much are revealing their sinophobia.
- adrian_b 2mo agoWhat I found the most interesting, and unfortunately not at all surprising, is that the reasoning of the LLMs frequently contained much more useful information than the actual answers, because the answers were censored.
- elzbardico 2mo agoDario is a cunning business man that won't hesitate to say whatever the fuck he needs to get the US government to exercise some regulatory capture to favor anthropic.
- dannyw 2mo agoYou don’t even get _any_ logits with closed models for years now. I can’t fault them too much, as logit based distillation is extremely effective. Very useful for making smaller models out of bigger open weight models.
- ziofill 2mo agoI understand it’s cool to have an artistic website, but it’s very noisy and non-accessible. But very interesting result.
- khalic 2mo agoThis is beautiful work, congrats
- SwellJoe 2mo ago"Stealing" is a strong word to use for looking at the words produced by models built from the collective commons of the world. And, honestly, being able to see how LLMs make decisions is critical to trust and security. I consider it a valuable feature, somewhat akin to seeing the source of software I use.
- dannyw 2mo agoYeah, it’s also useful for prompt tuning, debugging and understanding how a model interprets your prompt. Also really good for identifying any contradictions in your system prompt and context.
- happybox2016 2mo agoThe real issue is that API providers log everything. OpenAI/Anthropic already capture full CoT traces in their logs — they just don't expose them. Distillation via API is just making explicit what they already have.
- dboreham 2mo agoThe whole point of the encrypted payload returned to the client for future re-submission would be that they don't log.
- vinaigrette 2mo agoI must say right of the bat this is the best research paper/working paper in regards to its styling. Beautiful
- SwellJoe 2mo agoI agree on desktop/laptop, but on mobile there are images that appear under the text making it hard to read.
- user43928 2mo agoOn an iPhone Pro Max only the first trace is readable. Navigating to the right lands between two cards, so that neither is readable.
- deleted 2mo ago[deleted]
- thefourthchime 2mo agoI was going to comment on that. This is clearly a vibe-coded webpage. It sort of smells like GPT to me, or at least front-end design. But the author clearly went back and forth to make it beautiful. This is not the first output he got. This is the kind of stuff I point to when people talk about AI slop. AI is just a tool. You're still the person who has to deliver the output and have some taste.
- yetanotherjosh 2mo agoMy brain can't tell if the text is horizontal or slightly rotated. It's very hard to read. Beautiful to some, inaccessible to others.
- Cynddl 2mo ago> The providers did not acknowledge “any security implications arising from side channels or replay attacks.” All model providers acknowledged the receipt of our report and subsequently we were unable to launch the same attacks. I went straight to the ‘Responsible Disclosure’ section. Not surprising, but still disappointing.
- benob 2mo agoA natural next step is to use the reasoning traces to jailbreak the stronger models (https://arxiv.org/pdf/2603.12277 https://arxiv.org/pdf/2603.12277)
- certainforest 2mo ago+1
- tanh 2mo agoSo to make the APIs stateless (the "ideal" where they don't use server side sessions/etc) we ended up with this. I'm sorry but this is kind of hilarious. Given the salaries paid to the workers at these companies and the hype of the models, I can't believe they all fell to the same flaw.
- lukewarm707 2mo agostrange because, their subscriptions are not stateless. they log everything and send it to 3rd parties for moderation.
- redox99 2mo agoIt wouldn't matter if it was stored only on their servers. As long as they offer the feature to downgrade a chat to a dumber model that can be jailbroken (and the downgrade keeps the reasoning), this trick works.
- 8note 2mo agothis is a lethal trifecta, but where a chunk isn't even needed you have a secret to keep that is read by the llm, and untrusted input that wants to exfiltrate it. by hell or high water, the agent is gonna output that text
- varenc 2mo agoIf CoT wasn't stateless and you instead just got a reference which pointed to the CoT stored on the lab servers, the same vulnerability would still exist. Since you just need a weaker jailbroken model to read a smarter model's CoT. This being stateless or not doesn't really matter. The stateless part is also important for enterprise customers that require zero data retention. (they could scope CoT access per model, but then users couldn't switch models mid-session)
- dxsecarch 2mo ago[flagged]
- simonw 2mo agoThis is a neat attack against those encrypted reasoning blocks you get back from APIs like OpenAI and Gemini and Anthropic: > We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext. Should be easy for them to fix though: switch up the encryption key so it only works with the API for each specific model, rather than being shared across all of their models. And indeed, the paper says it's been fixed by all three providers (though no news on how they fixed it): > All model providers acknowledged the receipt of our report and subsequently we were unable to launch the same attacks.
- bonoboTP 2mo agoIt's not stealing.
- vhantz 2mo ago> For some AIME problems Opus 4.8 sometimes states the answer before deriving it. We find that the API summary does not always preserve this distinction, and can instead make the reasoning appear like a clean derivation. No surprise here but good to have more confirmation that they just put all that in the training data. And based on the "reasoning", the models have some form of index of those problems (or they are HEAVILY trained on them).
- throwa356262 2mo agoDidn't we see this with Fable 5 on multiple benchmarks?
- AbhinavX 2mo agoNot surprised. On many benchmarks (i.e tau), we have seen the same thing. Probably lots of training on every publicly available benchmark
- Aurornis 2mo agoAll LLM benchmarks have an expiration date once they're released to the public. They get spread so far and wide across the internet and GitHub that you have to assume they're in the training data for every LLM with a cutoff date after their release. The real question is whether or not the training was directed to optimize for those benchmarks. The technique doesn't guarantee that the reasoning is returned verbatim because it relies on the weaker model transcribing it accurately. Looking at the charts, there are a lot of dots that aren't in the 1:1 line that suggests that the output is exactly what was provided.
- tripzilch 2mo agoSo ... just a thought but could this be somewhat solved by, say if you were an LLM benchmark creator, using clever trickery? Like what if you made sure the wrong answers just appear 100x more often than the right ones. When scraping for new data to use I doubt they can verify the correctness of complex benchmark question answers to exclude the wrong ones. Then I dunno store the hash of the correct answers somewhere else, and eh try not to leak it. But even if it gets leaked, that just means perhaps at inference time, a clever agentic LLM could go for for those hashes and maybe determine what is correct, but not during training. I'm not sure, but wouldn't this make sure that at least they aren't literally trained on the correct question/answer pairs. I guess there would always be people that end up publishing the correct list, anyway. But that's why you try to be 100x "louder" with the wrong answers. btw, different thing, but when I look at those charts, I kind of came to the opposite conclusion as you did :) IMHO not that many dots off the line, and the ones that are on the line, are literally ON the line, not like a "roughly linear looking cloud of points". Which suggests that the reasoning is either (in the majority of cases) exactly the same amount of tokens (on the 1:1 line), and when it's even a little bit off the line it could (and should) be discarded, still leaving what seems to me at least 95% of the traces as exactly correct. but I grant, I didn't read the paper, and just came to that conclusion after viewing the chart :)
- unjuno 2mo ago[dead]
- elzbardico 2mo agoOpenAI and Anthropic will probably now resort to save this server side, instead of relying on encription to be able to keep state on the client.
- agenticfish 2mo agoThat's not a trivial thing to do for them because they offer zero data retention environments to enterprise clients.
- paxys 2mo agoThey can always keep the encrypted blobs server side and send the key to the user. But regardless, that isn’t going to help with this issue (see my comment above).
- paxys 2mo agoEncryption is irrelevant here. Even if it was kept fully server side, the actual issue is that they allow starting a conversation in a strong model and continuing it in a weaker one. Disallowing that entirely would be a huge hit to user experience.
- niemandhier 2mo agoYou cannot steal what is not owned. At least in the EU there is no copyright for LLM outputs, so I guess all they might do is violate the terms of service.
- Zambyte 2mo agoEven copyrighted information can never be "stolen". It can only copied without authorization.
- otterley 2mo agoStealing is not a word that applies only to physical objects.
- niemandhier 2mo agoFunnily enough in some legal systems it does. Where I live the legal definition of “theft” is: Taking away a movable thing.
- kube-system 2mo agoThat's also... a different word.
- niemandhier 2mo agoNot in my language.
- otterley 2mo agoWe're discussing the English language here.
- margalabargala 2mo agoAny word can be applied to any concept with any meaning thanks to the fluidity of vernacular. Language is all just sounds and markings. Anything can be redefined to mean anything, and anyone can decide to aggressively assert their preferred definition of a word.
- throwa356262 2mo agoThis is laughable security. People claim security is now "solved" thanks to AI but from where I am standings it looks more like the fun 90ies making a return. Anyway, can someone explain the part about K3? What are they trying to say?
- qrios 2mo agoThe interesting part is what they try to not to say: More indications for K3 is based on distillation from Claude and GPT. From [1]: > As you might guess, this suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography. > An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward Opus’s > A small memorization analysis showed that specific Claude and GPT reasoning spans are up to ~6 orders of magnitude easier to extract from Kimi-K3 than from the next-closest model. [1] https://x.com/kotekjedi_ml/status/2087147042888114428?s=42 https://x.com/kotekjedi_ml/status/2087147042888114428?s=42
- neuroelectron 2mo agoSecurity is solved, but business needs overrides it
- varenc 2mo agoThe part about K3 is just very strong evidence that K3 is partly a distillation of Opus. Probably even a distillation of Opus's CoT, which means they already had broken CoT encryption themselves awhile ago.
- deleted 2mo ago[deleted]
- andai 2mo agoIf I'm reading this right, they literally just ask a LLM to tell them what the traces say, with the key being that the traces are portable across LLM models, so they can switch to a smaller one that's easier to jailbreak.
- dgellow 2mo agoCorrect, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches
- HarHarVeryFunny 2mo agoFor all of the years of research, thinking and talking about model alignment, safety, confinement, etc, when it comes down to it these companies appear to be entirely incompetent.
- polymer8563 2mo agofor safety in particular it's pure theater, they only care as long as the orange guy thinks it's safe from "enemies of freedom"
- driverdan 2mo agoThis isn't a safety issue, it's LLM companies trying to be opaque and stop distillation.
- varenc 2mo agoBy some interpretations protecting your frontier model with strong safeguards from being distilled into an open source model without safeguards is a safety issue.
- realusername 2mo agoThere's no way to make a model "safe", (whatever that means) since you can't know what users will do with the output. It's just PR. Closed models are also used for nefarious usage.
- syntaxing 2mo agoPrefilling Kimi K3 with opus is a super interesting idea. That being said, I absolutely hate this website layout
- EagleEdge 2mo agoI used to do a very coarse version of this stealing. I ask a question from ChatGPT pro, once it is done, I ask claude chrome add-in to go through all those thinking from the side bar, extract everything along with all the sources used. Then try to reverse engineer the solution it came up with.
- arjie 2mo agoWow, almost certainly the approach that alternative labs use to distill Claude. I always wondered how far they could get with just the answer missing the reasoning. They probably actually also had the reasoning.
- C0ldSmi1e 2mo agoWhy they use different models to decode the reasoning content? Can the the model decode it?
- aszen 2mo agoBecause stronger models are harder to jailbreak from the paper it says haiku was easily fooled into giving us thinking contents
- sm-silversight 2mo agoIs this basically a paper on how to distill, in exactly the fashion openai/anthropic don't want/say is copyright theft?
- cush 2mo agoI really like this website
- pradeep1177 2mo agoThese logs containing opaque blobs could accidentally contain secrets, the researchers decoded many of reasoning blocks from public repositories and reported finding PII and credentials. I was experimenting a bit how I could block these using an ingress path. GitHub /softcane/hamza
- Pragmata 2mo agoApparently you can do the same by simply running it without reasoning, while giving it a thinking tool... >guys you do know you can just disable thinking, and instead give it a "deep_think" tool, and it will call it with internal CoT reasoning format right? >gl fixing that https://x.com/_can1357/status/2087228354399265125?s=20 https://x.com/_can1357/status/2087228354399265125?s=20
- retinaros 2mo agoits not exactly the same... its tool use spec asking to put thinking in inputs fields... it is a good idea but its not same.
- ashirviskas 2mo agoI've been doing that since before reasoning was a thing baked into the models, it always performs better this way. Except for some providers/models where you just can't easily turn it off, now I just avoid them. This way I save tokens and have full control of the reasoning.
- MaxMatti 2mo agoHow does it save tokens?
- klntsky 2mo agoThe model is finetuned to enter/leave its thinking mode using special token separators. there's no reason to assume the tool calls induce the same token distribution or produce the model's actual native reasoning trace
- its-summertime 2mo agoWe have its actual reasoning traces, and we have these psudotraces, distribution / nativeness is testable now
- sly010 2mo ago"Recovery" would be a more apt (although less catchy name). The stealing is on the provider side for not giving you access to tokens you already paid for.
- Havoc 2mo agoTIL it actually sends the traces. I had assumed this is entirely server side
- HoyaSaxa 2mo agoI can’t believe they don’t validate a decrypted signature belongs to the user or use a unique encryption key per user/session.
- Aissen 2mo ago"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge. Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economics-morally-charged-terms-and-distillation/ https://thomasdullien.github.io/posts/2026-06-15-rl-economic...
- deleted 2mo ago[deleted]
- __MatrixMan__ 2mo agoLiberating!
- paxys 2mo agoThe only person calling it stealing is the author of this article, so this is a pointless discussion. The majority of this thread is just arguing with themselves.
- dymk 2mo agoAnthropic and OpenAI made a big deal about how it's stealing.
- amazingman 2mo agoThey also have made a big deal about how what they did to build their models is not stealing. And we all know that's bullshit.
- blackqueeriroh 2mo agoNo, we actually don’t all know that.
- amazingman 2mo ago
- infecto 2mo agoWouldn’t the fix be to encrypt before the api call hits the LLM? You would encrypt/decrypt at a separate layer than the LLM. I am sure i am missing something but would love to be educated.
- paxys 2mo agoThe LLM needs to read the CoT as part of the conversation. You can ask the models to share them with you. Stronger models will refuse, while weaker ones can be “jailbroken”.
- infecto 2mo agoI don’t think that answers what I am wondering. Asked differently why is encoding/decoding the cot the concern of the LLM?
- paxys 2mo agoIt isn’t the concern of the LLM. Regardless of where the encryption/decryption is happening, the issue is that the LLM needs to access the raw CoT.
- infecto 2mo agoAgain I don’t think you’re really getting at what I am asking. Sorry. My whole point was why does the LLM have access of decrypting. It should happen outside of the LLM layer.
- bob1029 2mo agoI am slowly turning around on the idea of opaque reasoning tokens. In principle, yes, I want total control and visibility into the reasoning process. In practice, I find that it takes up so much time to DIY reasoning agents that I can't spend much energy on the actual application. The model providers have way more resources and talent to do this right and keep it right over time. I am willing to concede this moat to them if it means I can actually focus on the business. The more I think about it, the less I care to see those tokens. It feels an awful lot like obsession over logging every trace item an application could produce. Useful in theory but a nasty garage stacked to the ceiling with useless shit otherwise. What nefarious things are we concerned with? That they burn too many tokens in the black box? I'll threaten to move to a different black box. There are always options in a market with this many participants and big players feel this pressure. They know there's a point at which being evil bastards is no longer profitable. Spending time carefully designing tools and views that interact with the environment in clean ways is a much better investment than trying to own and control 100% of the reasoning process or models.
- retinaros 2mo agocurious seeing how anthropic is agressively fighting this stuff how did you get to experiment on this? did you just try and shown them results or did you need approval first? I am interested mostly because I research on distillation
- lossy_compress 2mo ago[dead]
- hahahaa 2mo agoYou wouldn't steal ... the token output you paid for.
- blmarket 2mo agoI expect future LLM will refuse to share the reason. "Hey, how did you come up with this idea?" "you have to pay enterprise API to learn this"
- glub 2mo agoI did this with Codex's recent encryption of compaction. Interestingly, I didn't have to drop to a dumber model, just a 2 sentence <developer> prompt auto-injected before and after compaction made all their models output the encrypted compaction data in plaintext. The result was... interesting. There's nothing unique in there and I still don't understand why they decided to encrypt it in the first place.
- chrisss395 2mo agoPossibly something to do with other providers using the it to train their own models?
- glub 2mo agoThe only "secret" there is a very basic instruction that the model receives, like "summarize current state and upcoming work" before compaction - same model that was just running your inference, with same cache, only server side, with no extra tools or capabilities. Then the fresh context gets the output from that as an encrypted blob + codex then injects up to 64k tokens of previous conversation, the latter part is visible in source code. There's nothing to gain from this, really. Perhaps they're preparing for something in the future, where they could give the model server-side tools that improves summarization, but right now, it's just a simple prompt.
- deleted 2mo ago[deleted]
- varenc 2mo agoThe compaction prompt doesn't seem like the valuable thing here. I suspect they're protecting the compaction result itself. If you're trying to distill a model, collecting lots of examples on how a large conversation gets compacted to a smaller summary is particularly useful data.
- glub 2mo agoNo, you can give the model same prompt and it will give you a similar compaction result. On the backend, that's precisely what happens. There's nothing else going on in that encrypted blob, it's just summary of what model responds with when prompted "summarize current state and upcoming work".
- tizerluo 2mo ago[flagged]
- varenc 2mo agosuper interesting. So pre-filling Kimi3 reasoning with Opus's reasoning results in thoughts that closely match Opus's. This seems like strong evidence Kimi3 was trained on decrypted Opus chain-of-thought. Meaning the Kimi team likely also broke CoT encryption. Though not exactly a big surprise.
- jijji 2mo agoThe fact that frontier LLM providers pirated all the data that they used for training, then to go on to encrypt all of the reasoning traces that they use to come up with the conclusions it's really disingenuous, and then have the balls to say distillation is some kind of bad behavior. they are the kings of distillation. The hiding of this data only brings distrust to their frontier models. I think most people want to understand how something comes to a conclusion they don't want have that part left out on purpose... it's this kind of behavior that forces people move to to open source models in the end, it's the lack of trust. the frontier model providers treat the end user/customer as a threat or adversary. Fable 5 is notorious for this. a lot of the serious questions you ask the model they won't even respond to you because of the woke guardrails. it wasn't only a couple weeks ago that huggingface had to use glm 5.2 to get the right answers about their security incident because Fable 5 didn't want to answer it.
- smeltworks 2mo ago[flagged]
- ggrab 2mo agoCool find, but can't help myself thinking that registering a domain name and submitting a paper on this to Arxiv is a bit... much. The content here could fit in a tweet or a short blog post as well. Not sure about the scientific novelty here as we're basically poking around the very top layers of someone else's software stack?
- Otterly99 2mo agoI really wonder how much of safeguarding with SOTA models is actually just "Don't do that" in a prompt? I'm building agent to control machines at work and have to rely on small models, which are especially bad at remembering rules like this, so it is very apparent to me that soft rules like this are pretty much useless. Might be different for these huge models, but still it seems very shaky to me.
- NegativeAbsence 2mo agoLast year's reports questioning whether reasoning blocks were actually reflected in the final response were why I stopped using reasoning models altogether. I switched to a separate pipeline and have used that ever since. It's good to see that the concern didn't remain just a suspicion.
- dylanw2468 2mo ago[flagged]
- aklein 2mo ago> Prefilling Kimi-K3's reasoning with the first 1% of tokens of Opus 4.8's reasoning moves its visible answer toward Opus's wording, even though the answer itself is never prefilled is this supportive evidence for the distillation accusations in the news?
- Fripplebubby 2mo agoPrefilling any model with the first 1% of reasoning tokens from another model should always move the output towards the output of the other model directionally - that's just next token prediction doing its thing.
- cryptonector 2mo agoA bit shocking. One would think that the encrypted reasoning traces -really, encrypted state cookies- would be bound to the session or user, not just the AI provider. That obviously is the fix.
- Bruce04 2mo agoI will repeatedly recommend RECOVERY DAREK to help anyone recover a lost coin or token. recoverydarek at Gmail dot COM helped me recover my stolen $162,000 worth of Bitcoin I Invested with a fake crypto trading company. All thanks to my cousin who referred me to Darek Recovery. In less than a week all my money was recovered.