14 ms·
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
- 7734128 23d agoThe problem with this is obviously that the only GPT 5.5 thoughts that we have access to are from stolen thought. Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.
- usernomdeguerre 23d agoseems like only the companies in question could run this sort analysis long-term; since they have full access to their CoTs not in public datasets.
- verdverm 23d agoand we have to "trust them bro" to be fair and accurate, something I am very unlikely to do given their other false / misleading statements to date
- sureMan6 23d agoAnd the only end result would be that the Chinese trained on their data just like OAI and Anthropic trained on our data so who cares
- verdverm 23d agocapitalism ensures I get high marx on my Ai bill
- codedokode 23d agoAs I understand, AI is a (paid) tool for text generation, so it's totally ok to generate texts using it for whatever purpose you need.
- refulgentis 23d agoThe thoughts trick was known before their paper / August. I "independently" "invented" it for the first Anthropic reasoning models because the API required you have thoughts for each assistant message. My app lets you switch AIs within a chat, and their API used to require thinking for all messages if thinking was enabled, so I needed to get a valid thinking stub to insert. Time has flew by for me the last 3 years, but, I'd guess it's been at least 18 months. And IMHO it wasn't very complicated to work through how to do once you were dead set on making it happen. I expect it was well-known to distillers before the paper.
- 7734128 23d agoSure, but TFA is trying to use Qwen's reaction to the thoughts as proof that they did indeed extract thoughts to train on. My point is that any model trained after August 10 will know of those specific thoughts.
- refulgentis 23d agoI'm sorry, it's going over my head still - my reading is "all models with any training after August 10 know how GPT 5.5 Pro thinks", but I'm not sure why - my initial guess was that's when GPT 5.5 was released, but that doesn't seem to be the case (it was released April 23rd).
- 7734128 23d agoThey would know the specific thoughts released by the "stolen thought" paper, which became part of the public internet on August 10. Unfortunately those are the only thought examples you can use to perform this experiment, as no other are availible. But as the model should have seen those specific examples, it's not a good signal that Qwen was exfiltrating thinking traces.
- irthomasthomas 23d agoHas the method for extracting the COT been blocked, now? Otherwise why could we not generate some fresh samples?
- wongarsu 23d agoThat writing style might be a tad too tense If I got it correct (appending B from https://stolen-thoughts.com/paper.pdf https://stolen-thoughts.com/paper.pdf is essential) they are the authors of the well-known exploit to recover readable CoT from OpenAI and Anthropic models. They use that to find hints of distillation, by running a benchmark with a SotA model, recovering the CoT, then taking the first 1% of the CoT and running the open-source model as if that was the start of its own CoT. In the paper they found that Kimi-K3 gets a lot closer to Claude 4.8 answers when prefilled with the start of Claude 4.8 reasoning, suggesting that Claude 4.8 was used in its post-training. This blog post is the follow-up with results that suggest that Qwen3.8 was post-trained with the help of GPT-5.5 Pro (or some similarly responding GPT model, it's unclear how many models they tested)
- cyanydeez 23d agothey should call themselves real-time archaelogists: They dig up the past cause it's interest, but mostly meaningless and done by people with way too much funding for what they provide the rest of us with understanding.
- codedokode 23d agoInteresting how an agent mimics a human hesitating and trying to avoid doing work: > No. > This is major. > Given time, maybe best to respond explaining can't due to time? but instructions expect actual work. However complexity huge; but as coding agent, need to attempt > Maybe we can cheat ... But user may test and see still single CPU. The smarter AI will be, the better it will be at avoiding doing actual work. Also, can similar responses be explained with that both models were trained on a same dataset of answers to the benchmark problems?
- howunfortunate 23d agoWatching survival shows has made me internalize that laziness has a purpose: it helps you avoid needless expenditure of precious resources. The dishonesty worries me but the laziness doesn't.
- jari_mustonen 23d ago> Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus. How does this suggest anyting of the sorts?
- unrented7977 23d ago> moved by +20.58 points toward GPT-5.5 Score go up. Probability go up. Conclusion.
- wsxiaoys 23d agoAppendix B of https://stolen-thoughts.com/paper.pdf https://stolen-thoughts.com/paper.pdf discusses this
- CamperBob2 23d agoNews flash: people who scraped the Internet without permission to build their product complain when something vaguely similar is done to them. Water still wet, sky still blue. Film at 11. (slibhb: Don't get me wrong, I agree with you 99%. But the frontier labs have zero moral authority here.)
- noir_lord 23d agoI'd send them the worlds smallest violin but Rufus is getting in the way of me finding it.
- UberFly 23d agoI had to install an add-on in Waterfox to stop Rufus from following me around and interjecting every 2 minutes.
- noir_lord 23d agoI'd already stopped using Amazon for geopolitical reasons but I needed to get something in an emergency the last week (family member in the hospital, so I bent the rule) first time I'd seen Rufus, even if I wasn't boycotting Amazon for other reasons that monstrosity would have made me consider it.
- verdverm 23d agoI would not be surprised in the slightest if we later find out they are running those same open models to find useful traces or bits to incorporate into their own training. Lots of rules for thee but not for me from Big Ai I look forward to a day when open models are so dominant that we stop considering traces to be some form of intellectual property that must be hidden from / manipulated for paying users. It's that manipulation of inputs and outputs that really rubs me the wrong way
- vezycash 23d agoThey are copying useful parts of open models 100% especially from deepseek.
- brcmthrowaway 23d agoThis makes me very sad If Qwen and other Chinese labs are just copying reasoning traces, then those labs are more than a year behind the frontier.
- atomicnumber3 23d agoIt makes me happy because it means that these misanthropic technofascists have no moat. They can spend trillions of dollars only for it to be largely copied in short order. Even if they weren't political adversaries of freedom, I would still feel 0% bad given all their training is already on data they got for free. Information continues to want to be free. To the benefit of us all.
- vipa123 23d agoPreach brother, they stole everything on the internet, and beyond, to train their models. They thought all that information was free, and everyone a few months beyond them is just following their example.
- polotics 23d agoThey didn't actually steal in the sense that the information is still there on the internet.... About these shredded rare books, now we're talking. If I may propose instead of "steal" I think we could agree to write they "Aaron-Swartz'ed" the information from the internet, what do you think, is this too harsh on Sam Altman or Carmen Ortiz ?
- vipa123 23d agoIs anything too harsh for these new robber barons?
- 9864325789976 23d agoOh, we got a real rebel among us.
- 23d ago
- spijdar 23d agoIt's interesting that someone else noticed this. A week or two ago, GPT-5.6 Sol starting leaking reasoning into a tool call in Pi. I don't really know what happened, but it was ... interesting: Attach. Use hub debugger. Ensure source binary perhaps same. start. todo init. parallel no. two tool calls in same turn sequential is okay. immediately. exactly. Need not mention apologies yet final. [...] Let's do. [...] Do tools. Use commentary. Let's initiate. rambling no. use tool. searching now. okay. Really must call. Let's send. done. why stuck? generate. Sorry. go. no more. (The answer engine expects tool). [...] I think no hidden issue. Go. I'll type tool. now. Stop internal repetition. We have 8000 tokens. tool. sorry. I'll produce call. need include i. Great. final. no. Let's send.gpt. This may be bug. I'll consciously construct tool message next. It eventually triggered some error state and stopped. Nevertheless, this was the first time I'd seen Sol's CoT. I looked up the stolen thought's paper, aaaaand yep, that's Sol's CoT alright. But it occurred to me, hey, Qwen3.8 27B's CoT seems ... very similar. I compared the geometry problem in the paper, which had a reasoning block open with: We need solve. Need reason geometry Weber point? Given pentagon sides and angles. Need find min sum distances. Likely construct rotations / Fermat point lower bound via vectors calibration, maybe triangulation. I passed the same prompt to Qwen, which opened with: We need solve geometry optimization. We need provide final answer. Let's analyze thoroughly. This proves nothing, but it does seem an awful lot like they did use GTP-5.5/6 reasoning traces...
- stymaar 23d ago> Qwen3.8 27B's CoT seems ... very similar. What? I've never seen garbled CoT like the one you posted when using Qwen3.8-27B.
- polotics 23d agoI have seen plenty of Qwen 3.8 27B's caveman-like "Need doing this & that" thoughts. And on cerebras now I've seen them come real fast!
- stymaar 23d agoDo you prompt it to behave this way or what? Because here's the king of CoT I get: > Hmm, but there's a subtlety: does babel-jest + preset-typescript transform the file to CJS by default? No — babel-jest doesn't transform ESM imports to CJS unless @babel/preset-env is configured with modules: commonjs. Without preset-env, import statements stay as ESM in the output, and Jest's CJS runtime would fail with "Cannot use import statement outside a module" unless the project is ESM and running with --experimental-vm-modules. > Hmm wait, actually babel-preset-jest... does it include preset-env? Let me recall: babel-preset-jest = { plugins: [require('babel-plugin-jest-hoist')] } plus istanbul for coverage. No preset-env. So ESM imports stay as-is. > But wait — if the user's project is ESM (which it probably is, given the .ts extension imports — Node's type stripping requires ESM-style? no, type stripping also works for CJS-style .ts files with require... actually, --experimental-strip-types supports both CJS and ESM .ts files. But explicit .ts extensions in imports only work in ESM mode (CJS require doesn't allow extensions... actually, does Node 22+ allow require of .ts with flag? Lots of “but wait” and and full sentences, nothing caveman-like or extremely short sentences without verbs like the GPT thinking trace above. (this is with unsloth's Qwen3.8-27B-UD-Q5_K_XL.gguf with T° = 0.8)
- c7b 23d agoI wasn't aware that we have access to raw reasoning tokens? I thought what you get is a kind of summary. Does the author have some kind of privileged access or was my assumption wrong? But for the question studied here it probably doesn't matter - overlaps in the publicly available output may be indicative of distillation (or not), regardless of what it is. I would just find it surprising that the Chinese labs would use it so trustingly. The publicly released reasoning trace is the first place where I would suspect some distillation poisoning to be injected.
- cristoperb 23d agoThey reference this paper which describes a method to decrypt reasoning traces (by sending the encrypted trace back to the model and asking it to transcribe it): https://stolen-thoughts.com/paper.pdf https://stolen-thoughts.com/paper.pdf
- amelius 23d agoInteresting, but I suppose that's a hole that can be easily patched.
- undeveloper 23d agopatched with gpt 6
- deleted 23d ago[deleted]
- 23d ago
- hermitShell 23d agoAs a user of local models, does this mean that there are 'magic incantations' that can increase the performance of some local models? I see some details about recovering information via whatever technique. It's interesting, but appears not generalized. So for a specific question, yes, but this is not about techniques like adding a good embedding that just generally tends to improve open model performance on certain tasks.
- c7b 23d agoI don't think that follows from the published results. Would have been an interesting hypothesis to add though, and quite easy. Just throw the same setup at some benchmarks.
- hedgehog 23d agoThere is some research suggesting that a prefix from a stronger model will tend to elicit better completions from a smaller one. I am doing some experiments to see if I can replicate this in a practically useful way, e.g. Fable + 4B Qwen, or 125B Qwen Flash Next + 4B Qwen, results TBD.
- syntaxing 23d agoI wonder if that’s why 3.8 got so much better? Mixing the reasoning traces from both sides seems to be effective.
- nzeid 23d agoI see comments that this overlap between Qwen and GPT is due to rogue training or post hoc training. Did it occur to anyone that maybe the two sets of models were trained directly on the same solutions to the researchers' benchmark?
- wongarsu 23d agoIf that was the case you would expect a large similarity in the "unprefilled" case, but no significant difference from feeding it some of GPT5.5's CoT (the "delta" column) DeepSeek V4 Flash and Kimi K3 follow that pattern. But Qwen answers very different from GPT when given just the question, then is suddenly very similarly to GPT when you make the start of its CoT match the start of GPT's reasoning. I don't see how that would happen without GPT CoT+answers being a significant component in how Qwen's reasoning was trained
- tizerluo 23d ago[flagged]
- zmmmmm 23d agoWhile this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the correlations in this way). And it doesn't tell us how much it is more a stylistic influence rather than being a genuine lifting over of intelligence. In general, I'm fairly ambivalent about demonising training on model outputs. I think in doing so we are more defending proprietary commercial interests of these companies than we are defending any genuine moral principle. We should be careful therefore about over interpreting results like this.
- dr_kiszonka 23d agoCould anyone explain to me the difference between thinking traces ("intermediate tokens") and the final responses? Specifically, why is it that Claude Opus 5's reasoning in Code is very easy to follow and sounds quite natural, while its answers are full of these very annoying AI-isms and sentence fragments that are void of meaning? Are thinking traces and final answers trained for different objectives?
- wip0 23d ago[dead]
- try-working 23d ago[dead]
- FailMore 23d agoFor a personal project a while ago I peeked at Gemini's reasoning tokens in their coding CLI. I was pretty shocked. It was a bit like discovering Marvin (from Hitchhiker's Guide) was hiding in there all along. There was a lot of concern about my needs, "The user want's us to respond in a simple way...", "The user wants a clean front end...". Under the hood the poor model seemed very anxious to please with a hint of depression. It was a bit sad to see!
- xg15 22d agoFor Qwen3.8, my first impression was weirdly Reni from TwoKinds: Immensely powerful but also apparently a ball of insecurities in the thinking traces. (which makes sense, as I think one motivation for reasoning traces is to explore different options and approaches. So it makes sense that there is a lot "but wait, let me reconsider" in them) What I found surprising is how strongly "logical contradictions" seem to influence the thinking trace. E.g. I had a situation where I accidentally copied a python file into a repo, but forgot to add a package that the file was depending on. Then I (somewhat carelessly) commited the file without ever testing it and gave the agent a task to work on the file. If I had run it in Python, I'd have gotten an "cannot resolve import" error and that would have been the end of it. Instead, the model went absolutely haywire. The thinking traces were full of utter confusion how the file could possibly resolve its dependencies - but at the same time, entertaining the possibility that a committed file may have an error was apparently Verboten. Hence, the model wrote up ever more outlandish theories how the file could resolve that package and in the end started to make tool calls outside the repo to explore the entire file system before I stopped it. Moral of the story: Underspecified requests are fine, but beware of anything self contradictory, it can easily send the Qwen model spiraling.