8 ms·
What GPT-OSS leaks about OpenAI's training data
- zaptrem 1y ago> There are about 936 tokens with very low L2 norm, centered at about 2. This likely means that they did not occur in the training process of GPT-oss and were thus depressed by some form of weight decay. Afaik embedding and norm params are excluded from weight decay as standard practice. Is this no longer true? E.g., they exclude them in minGPT: https://github.com/karpathy/minGPT/blob/37baab71b9abea1b76ab957409a1cc2fbfba8a26/mingpt/model.py#L219 https://github.com/karpathy/minGPT/blob/37baab71b9abea1b76ab...
- 3abiton 1y agoUnfortunately the article glances over some of practices of uncovering such patterns in the training data. It goes very straitghfully to the point, no lube needed. It didn't land well for me.
- levocardia 1y agoCould it instead be the case that these tokens were initialized at some mean value across the dataset (plus a little noise), and then never changed because they were never seen in training? Not sure if that is state of the art anymore but e.g. in Karpathy's videos he uses a trick like this to avoid the "sharp hockey stick" drop in loss in the early gradient descent steps, which can result in undesirably big weight updates.
- behnamoh 1y agoIs there any work on reverse engineering LLMs, especially the closed source API ones? For example, how can we learn about the data used in Claude Sonnet 4.5 training? And more tricky but as important, is there any work on extrapolating the pretrained model AFTER it's RLHF'd? For example, what kinds of biases did exist in gpt-4o before it was unbiased? Do biases go away completely or they just get suppressed down deep in the model's "mind"?
- tptacek 1y agoYes. https://arxiv.org/abs/2403.06634 https://arxiv.org/abs/2403.06634 https://arxiv.org/abs/2311.17035 https://arxiv.org/abs/2311.17035 (I just have these ones off the top of my head because I'm a Nicholas Carlini fan and we interviewed him about these attacks.)
- behnamoh 1y agoThanks for these, I'll have a look!
- zer00eyz 1y ago> Do biases go away completely or they just get suppressed down deep in the model's "mind"? Bias is a human term, and couching the conversation in that context does nothing to address the issue here, because it gets into the quagmire of social context. Let's say LLM's had taken off 15 years ago at the point system d launched. All the answers given are going to weight toward the old init system simply because there is a lack of information. LLM's are only repeating the data they are given, and it's cheaper to remove the data after the fact than it is to try to scrub it out of the training data.
- astrange 1y ago"only" and "repeating" aren't accurate here. There's a lot of steps between the pretraining tokens and the LLM. I mean, you can pretty much do whatever you want in the process of making one or running it. For instance you could use pretraining/SFT to steer something away from a document instead of towards it and that wouldn't be "only repeating" it. Though I don't know if that's actually possible, and afaik it is true RL reweights existing data instead of learning new things.
- lupusreal 1y ago> Bias is a human term There are many kinds of bias, plenty of which have nothing to do with culture or social context.
- Wowfunhappy 1y agoMaybe I'm misinterpreting, but the article seems (?) to be implying there's something scandalous about OpenAI training an adult websites. I find that odd. Would anyone be surprised to know that Google indexes adult websites, and ranks them in its search algorithm? If not, what is the difference for an LLM?
- refulgentis 1y agoFWIW, I didn't get that sense.
- raincole 1y agoAnd it's nothing new. https://github.com/jiangyy/gpt-tokens https://github.com/jiangyy/gpt-tokens People found these adult-site-related Chinese phrases in Gpt-4o. The OP is more than one year late.
- pydry 1y agoTheyre saying if you find references to a very specific set of phrases that were probably included accidentally on github then github is likely part of the training data.
- relatedtitle 1y agoGitHub is obviously part of the training data, you don't need to find obscure tokens to tell.
- mudkipdev 1y agoWouldn't it be best for them to strip that out of the training data for moderation reasons?
- rs186 1y agoMany of the crude translations of those Chinese phrases are way off to the point that it fails to understand the meaning, which makes me think the data in those matrices is inaccurate as well. The author really needs to ask a native Chinese speaker with experience in ... searching explicit content to proofread the article and examine the results.
- Theodores 1y agoFascinating article. I am giving everything AI a wide birth for now, however, I do enjoy learning about how AI works. The question I have, is what does a LLM do when it encounters a new token? Can it actually learn from context, etymology and usage? As I child I had no idea what many of the words meant in the newspaper and in literature but I could just pretend I knew what those words meant or get by without knowing what those words meant in full. In time I would gain familiarity with these words, able to make sense of them in context but not necessarily able to pronounce said words or be able to use them in my own writing. I certainly didn't stop what I was reading to get the dictionary out every time I encountered a new word, and this is how I think most people learn to read, with gradual changes with new words going from no idea to some familiarity to confidently able to use. We aren't tokenising like the LLMs do and our languages are the product of many hundreds of thousands of years of development. So, how does an LLM learn words that have not already been tokenised? Or is this baked in?
- martin4114 1y ago[dead]
- refulgentis 1y agos/birth/berth :)
- DrewADesign 1y agoThat's rather presumptuous, don't you think? There are some people here with very unusual jobs.
- martin4114 1y ago[dead]
- martin4114 1y ago[dead]
- FeepingCreature 1y agoInformed layman warning. The tokenizer covers the entire dataset. It's basically just a fixed-size Huffman code, grouping together common fragments of letters- for instance, the 100 most common English words are probably all single tokens. During learning, the model proceeds in roughly the same way a child would: it starts by grouping tokens together, learning the deep regularities of language such as "news[paper]" being more likely than "news[q77.bfe]". Then it incrementally assembles these fragments into larger and larger chains. Similarly , it first learns thematic groupings, such as "word" being more likely somewhere after "dictionary" rather than "stop what I was reading to get the dictionary out every time I encountered a banana assault hungry". Then it starts to pick up "patterns": "as a [baby|child|kid] I had no [idea|concept|clue]". At some point in this process it naturally abstracts concepts from languages: "as a child" starts being internally represented by the same neurons as "als ich ein Kind war". Then some magic happens that we don't understand, and out pops a neural network that you can talk to and that can write programs and use tools. To be clear, this is the case before RL: probably these patterns are now widespread in the training data, so that the model already understands how to "complete the pattern" on its own. RL then does some magic on top of that to bring it from 20% benchmarks to 80% and presto, AI assistant.
- httpsoverdns 1y agoI tried many of the examples in this article in Gemini 2.5 pro and it seems to handle most quite flawlessly. Is it possibly that Google's model is just susceptible to different glitch tokens? I admit most of the technical discussion in the article went a little over my head.
- simonw 1y agoGlitch tokens should be tokenizer-specific. Gemini uses a different tokenizer from the OpenAI models. The origins of the OpenAI glitch tokens are pretty interesting: the trained an early tokenizer on common strings in their early training data but it turns out popular subreddits caused some weird tokens to be common enough to get assigned an integer, like davidjl - a frequent poster in the https://reddit.com/r/counting https://reddit.com/r/counting subreddit. More on that here: https://simonwillison.net/2023/Jun/8/gpt-tokenizers/#glitch-tokens https://simonwillison.net/2023/Jun/8/gpt-tokenizers/#glitch-...
- magicalhippo 1y agoGiven that the token space is large enough to waste on such "low quality" tokens, has there been work done to use a smaller token space in order for quantized models to perform better? Just a silly thought that crossed my mind when I saw those "ad tokens".
- typpilol 1y agoIsn't that exactly what some of these models that have 30b params but only activate 3b at a time
- rvba 1y agoHumans also only use X% of their brains (the one needed for a specific task)
- koakuma-chan 1y agoDoes that mean I'm a mixture of experts.
- RLAIF 1y ago[dead]
- clarionbell 1y agoThat's mixture of experts pattern.
- magicalhippo 1y agoAs mentioned that would be more like mixture-of-experts, where you instead effectively just multiply the rows with some of the columns in your weight matrices. That said, I wrote my post after bed time and I'm now pretty sure I wasn't thinking straight.
- NoahZuniga 1y agoThis article says that "GPT-5 was trained on phrases from adult websites". However, this is misleading as the only thing that was shown is that GPT-5 was trained on phrases that also occur on adult websites, with some speculation of the source of the training data container such adult phrases being GitHub.
- tymscar 1y agoThis is addressed at the end of the blogpost
- a_victorp 1y agoIt is not
- rahulstein 1y agoIt is - in the link to the MIT Technology Review article
- jimmydoe 1y agoChinese adult site ads are everywhere in repackaged free and pirate content, which are distributed thru sites including but not limited to github, shadow libraries and youtube. for same reason, whisper some blank audio will output those ads.
- breakingcups 1y agoSpecifically, because some pirates will put advertisements for other illicit services at the beginning or end of movies and tv shows in its subtitle data where there's a suitable gap. Usually those gaps have silence. Companies incorporating subtitle data as transcription source of truth training data will thus train their models to output facsimiles of these messages whenever they're encountering prolonged stretches of silence.
- wongarsu 1y agoI think that is what the article is saying, just with a very misleading phrasing. The phrases are from adult websites, and likely entered the training data via Github (lists for content blockers etc). Just like "to be or not to be" is from Hamlet, but you just read it on HN. At no point does the article actually say that adult websites were in the training data
- starkeeper 1y agoI wish we had a constitutional amendment that opensourced all AI commercial AI models and requires documentation and links to all training data and base prompts. They are trained on public data at our expense so We The People should *own* them. Someday probably sooner then we might even think.... We'll easily run mega huge sized models on our laptops, desktops, and phones. AI should be free. Overhyped and Overpriced. I would love this setup for privacy and security. Anyways, only tangentally related... (why worry about leaks like this and the hidden base prompts! - they *should all be 100% OSS* - it is the only way to ensure privacy and security). Also, long timer lurker, first time posting! I just had to get this off my mind! Cheers.
- heavyset_go 1y agoI'd settle with them being held in a public trust for public benefit
- halperter 1y agoUnfortunately very unlikely in our forseeable future with the U.S. having a "U.S. against the world" mentality to the AI race. Would love to see this but this would get shot down immediately.
- canadiantim 1y agoWouldn’t the same argument then be applied to all scraped data?
- rileymat2 1y agoWhy would it require a constitutional amendment?
- delichon 1y agoThe takings clause of the fifth amendment allows seizure of private property for public use so long as it provides just compensation. So the necessary amendment already exists if they're willing to pay for it. Otherwise they'd need an amendment to circumvent the fifth amendment, to the extent the document is honored.
- renewiltord 1y agoInteresting. Small typo by the way. It's SolidGoldMagikarp with a k. Easy mistake to make with that tokenizer though har har It strikes me less that they're from adult websites and more that they're from compromised sites. I've had that happen before and it's mostly porn and stuff like that when that happens.
- smj-edison 1y agoOne interesting tidbit from this article that I haven't seen mentioned yet is that you can use glitch tokens to figure out what model someone is using behind the scenes. Put a glitch token in a prompt, and see if it reacts normally or response with this kind of glitchy behavior.
- sails 1y agoYes I thought so too. I wonder if it will mean more or less revealing of the models that are running agentic flows (we currently abstract with Fast/Smart) It is also possible that the first model calls other models, and you could reverse engineer the tool call structure by seeing when glitches occur based on different branches of the tool calling
- willvarfar 1y agoYou can imagine LLM fingerprinting to be part of future pentest workflows where they identify the model and know it's weaknesses and vulnerabilities etc...
- lyu07282 1y agoIsn't the only reason we can do that because we have access to the tokenizer? Do we have the Claude and Gemini tokens? I mean if they didn't publish that, would it defeat this attack?
- sebzim4500 1y ago
- indrora 1y agoThere's an interesting set of options for the weird "xadder" token: misspellings of "xpadder" (a game pad helper), xadder (the name of at least two or three tools), xadder (a parameter in an XLib call), XAdder (the Xilinx full adder implementation for the Vivado FPGA platform), and more than a few usernames on various forums.
- blochist 1y agoImportantly, though, the token is "\\xadder" which looks a bit like an escaped hex code. That actually suggests a different origin of the token. `\xad` is the Unicode soft-hyphen (U+00AD). The soft hyphen is used to suggest where it makes sense to hyphenate a word if a line-break is needed. This shows up fairly frequently (2.9k occurrences) on GitHub in web-scraping datasets, which suggests that a model trained on data scraped from the web might see a fair number of these. Basically, OpenAI is trained on web data that has a number of words where splitting on -der makes sense (e.g., mur-der, un-derstanding, won-derful; although the most common occurrence in a GitHub search for "\\xadder" is what appears to be an incorrectly encoded string "L\xc3\xadder", probably from the Portuguese and Spanish "Lí-der"). Anyways, using the o200k tokenizer `mur\xadder` yields two tokens (88762 and 179582). 88762 encodes "mur" and 179582 encodes "\xadder".
- supermatt 1y ago> GPT-5 was trained on phrases from adult websites Does it really imply they were trained on phrases FROM adult websites, or that those phrases FOR adult sites were common in the training data? Blogspam, link-farms, affiliate marketing, etc, are extremely common for adult (and gambling) sites and likely result in a lot of data tainted with those phrases.
- dbtablesorrows 1y agoThis guy adults.
- DonHopkins 1y ago>All of OpenAI's models since GPT-4o use the o200k tokenizer. I wonder what the full 202 letter name of the o200k tokenizer is?