5 ms·
Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obses
by postalcoder 5mo ago
Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand:
- The sepia tint on images from gpt-image-1
- The obsession with the word "seam" as it pertains to coding
Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people saying that frequently.
- operatingthetan 5mo agoSeams, spirals, codexes, recursion, glyphs, resonance, the list goes on and on.
- krackers 5mo ago>with the word "seam" as it pertains to coding I thought this was an established term when it comes to working with codebases comprised of multiple interacting parts. https://softwareengineering.stackexchange.com/questions/132563/problem-with-understanding-seam-word https://softwareengineering.stackexchange.com/questions/1325...
- postalcoder 5mo agothanks for this. > the term originates from Michael Feathers Working Effectively with Legacy Code I haven’t read the book but, taking the title and Amazon reviews at face value, I feel like this embodies Codex’s coding style as a whole. It treats all code like legacy code.
- eterm 5mo agoIt's been a long time since I read it, but it was one of the better books I've read. It changed my approach to how to think about old code-bases.
- kqr 5mo agoI agree. I come back to it all the time when I need a little inspiration for how to deal with a gnarly codebase. Usually there is something in there I can apply directly to get me out of a pinch. When there is not the reminder of how malleable code is suffices.
- TeMPOraL 5mo agoIt's not in the top 10, but it's of the more well-known and widely recommended book in the software industry. I'd put it in the same bucket as "Clean Code" and maybe even "Domain Driven Design"; they're kinda from the same "thought school" in the software industry. So it's definitely over-represented in training data (I'd guess primarily in the form of articles and blog posts and educational material reiterating or rephrasing ideas from the book). FWIW, I found the concept of "seams" from that book useful back when working on some legacy C++ monolithic code few years back, as TDD is a little more tricky than usual due to peculiarities of the language (and in particular its build model), and there it actually makes sense to know of different kind of "seams" and what they should vs. shouldn't be used for.
- tdeck 5mo agoI can't say it isn't, but I have been writing code since about 2004 and this is the first time I've become aware that this is a thing.
- jofzar 5mo agoOne I saw recently was "wires" and "wired" from opus. It was using it like every 3rd sentence and I was like, yeah I have seen people say wired like this but not really for how it was using it in every sentence.
- baq 5mo agoGPT started to ‘wire in’ stuff around 5.2 or 5.3 and clearly Opus, ahem, picked it up. I remember being a tiny bit shocked when I saw ‘wired’ for the first time in an Anthropic model.
- Barbing 5mo agoAnthropic distills GPT?
- yorwba 5mo agoEverybody training models on large amounts of lightly filtered internet text is partially distilling every other model that had its output posted verbatim to the internet.
- beAbU 5mo agoAnd OpenAI probably distills anthropic, who would't? It's all one big incestuous mess. In a couple of years we'll be talking about AI brainrot.
- NitpickLawyer 5mo agoAll GPTisms are like that. In moderation there's nothing wrong with any of them. But you start noticing them because a lot of people use these things, and c/p the responses verbatim (or now use claws, I guess). So they stand out. I don't think it's training data overrepresentation, at least not alone. RLHF and more broadly "alignment" is probably more impactful here. Likely combined with the fact that most people prompt them very briefly, so the models "default" to whatever it was most straight-forward to get a good score. I've heard plenty of "the system still had some gremlins, but we decided to launch anyway", but not from tens of thousands of people at the same time. That's "the catch", IMO.
- pants2 5mo agoMaybe the only solution to GPTisms is infinite context. If I'm talking to my coworker every day I would consciously recognize when I already used a metaphor recently and switch it up. However if my memory got reset every hour, I certainly might tell the same story or use the same metaphor over and over.
- telotortium 5mo ago> However if my memory got reset every hour, I certainly might tell the same story or use the same metaphor over and over. All people repeat the same stories and phraseology to some extent, and some people are as bad or worse than LLM chat bots in their predictability. I wonder if the latter have weak long-term memory on the scale of months to years, even if they remember things well from decades ago.
- yard2010 5mo agoHonestly I think there is more to it - even with infinite context, the LLM needs some kind of intelligence to know what is noise and what is not, you resort to "thinking" - making it create garbage it then feeds to itself. Learning a language is a big complex task, but it is far from real intelligence.
- yard2010 5mo agoI think the problem is that humans are not random, they are very biased. When you try to capture this bias with an LLM you get a biased pseudo random model
- vunderba 5mo agoIt was always funny how easy it was to spot the people using a Studio Ghibli style generated avatar for their Discord or Slack profile, just from that yellow tinging. A simple LUT or tone-mapping adjustment in Krita/Photoshop/etc. would have dramatically reduced it. The worst was you could tell when someone had kept feeding the same image back into chatgpt to make incremental edits in a loop. The yellow filter would seemingly stack until the final result was absolutely drenched in that sickly yellow pallor, made any photorealistic humans look like they were all suffering from advanced stages of jaundice.
- andai 5mo agoFor context, an example of what happens when you feed the same image back in repeatedly: https://www.instagram.com/reels/DJFG6EDhIHs/ https://www.instagram.com/reels/DJFG6EDhIHs/
- vunderba 5mo agoHaha fantastic. I'd love to see a comparison reel of that same image-loop for the entire image gen series (gpt-image-1, gpt-image-1.5, gpt-image-2).
- dmichulke 5mo agoFixed points are a window to the soul of a LLM - Lucretius in "De rerum natura", probably
- Suppafly 5mo agoI like how the AI seems forced to change their ethnicity to keep up with the color changes. Absolutely wild.
- Barbing 5mo agoMirror: https://files.catbox.moe/mu8env.mp4 https://files.catbox.moe/mu8env.mp4
- 5mo ago
- alex_sf 5mo ago"shape" too, at least with gpt5.5, is coming up constantly.
- tudorpavel 5mo agoThe one phrase that irks me as overly dramatic and both GPT and Claude use it a lot is "__ is the real smoking gun!" I'm a non-native English speaker, so maybe it's a really common idiom to use when debugging?
- aorloff 5mo agoIt probably was found in a bunch of meaningful code commit messages
- socks 5mo agoMy colleagues were joking about smoking guns yesterday after noticing that Claude was obsessed with it.
- thinkingemote 5mo agoI like how your co-workers enjoy the language. I had a similar group of colleagues once who did similar pre LLM but with words in popular culture, very playful. In the future these tells will be more identifiable. We will be easier to point back at text and code written in 2026 and more confidently say "this was written by an LLM". It takes time for patterns to form and takes time for it to be noticeable. "Smoking gun was so early 2026 claude".I find thinking of the future looking at now to be refreshing perspective on our usage.
- gizajob 5mo agoI’m a British English speaker and find the use of cliched American idioms really quite disgusting. Don’t want to think about about ballparks, home runs, smoking guns, going all in, touchdowns or hitting it out the park.
- weitendorf 5mo agoIt actually probably wouldn’t be too expensive or difficult to finetune those sayings out of default behavior if it were made accessible to you, you could even automate most of the relabeling by having the model come up with a list of idioms and appropriate replacement terms so it calls eg cookies biscuits or removes references to baseball. Absolute bollocks they don’t offer that as a simple option anymore
- pdntspa 5mo agoThe number of things that Claude has told me are 'load-bearing' or 'belt-and-suspenders' is... very load-bearing
- DespairYeMighty 5mo agofor me, doing the heavy lifting is doing the heavy lifting
- andromaton 5mo agoAlso too many lands and hits.
- yard2010 5mo agoFun fact: the word suffer comes from sub fer - under load, this relation (suffer - load bearing) is consistent across (unrelated) languages
- sushid 5mo agoYou are absolutely right to call that out!
- vidarh 5mo agoClaude, at least 4.5, not checked recently, has/had an obsession with the number 47 (or numbers containing 47). Ask it to pick a random time or number, or write prose containing numbers, and the bias was crazy. Also "something shifted" or "cracked".
- dhosek 5mo agoHumans tend to be biased towards 47 as well. It’s almost halfway between 1 and 100 and prime so you’ll find people picking it when they have to choose a random number. Then there’s the whole Pomona College thing https://en.wikipedia.org/wiki/47_(number) https://en.wikipedia.org/wiki/47_(number)
- vidarh 5mo agoThe whole blue 7 thing [1] and variations is very fascinating, but we don't tend to repeatedly pick the same number in the same exact context, though. That's what made this stand out to me - I had a document where Claude had picked 47 for "random" things dozens of times. [1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon I experienced this even second hand when a coworker excitedly told of an encounter with a cold reader, and I knew the answer would be blue 7 before he told me what his guess was. Just his recap of the conversation was enough.
- flawn 5mo agoI am biased towards 67
- eterm 5mo ago"is the real" is such a strong Claude tell, whenever I encounter it, it makes me question what i'm reading. Another I've noticed more recently is a slight obsession over refering to "Framing".
- ahmadyan 5mo agoi just want to know where emdash came from, as it is quite rare to see it on the public internet, so it must have been synthetically added to the dataset.
- LiamPowell 5mo agoThe very simplified answer is that the models are first trained on everything and then are later trained more heavily on golden samples with perfect grammar, spelling, etc..
- doginasuit 5mo agoEmdash is very common in academic journals and professional writing. I remember my English professor in the early 2000s encouraging us to use it, it has a unique role in interrupting a sentence. Thoughtfully used, it conveys a little more editorial effort, since there is no dedicated key on the keyboard. It was disappointing to see it become associated with AI output.
- deleted 5mo ago[deleted]
- TeMPOraL 5mo agoOther than things other comments already mention, let's not forget that Microsoft Word auto-corrects "--" to em-dash, and so does (apparently - haven't checked myself) Outlook, Apple Pages, Notes and Mail. There's probably bunch of other such software (I vaguely recall Wordpress doing annoying auto-typography on me, some 15 years ago or so).
- gizajob 5mo agoBecause on the public internet people don’t have arts degrees which are where emdash users learn to wield it correctly.
- dboreham 5mo agoI learned about em-dashes by reading Knuth about 40 years ago.
- dyauspitr 5mo ago“I’ve got the shape of it now”
- croisillon 5mo agoand "quietly"!
- wodenokoto 5mo agoI thought the “why it matters” headline was a funny reference to ChatGPT phraseology
- teaearlgraycold 5mo agoWhenever Claude finishes some work it almost always says “Clean.” before finishing its closing remarks. It’s at the point where I repeat it out loud along with Claude to highlight the absurdity of the repetition.
- weitendorf 5mo agoWith 4.5, I think because I would prompt it/guide it towards an outcome by calling it “the dream: <code example>” it would get almost reverential / shocked with awe as it got closer to getting it working or when it finally passed for the first time. Which was funny and reasonably context appropriate but sometimes felt so over the top that I couldn’t tell if it also “liked” the project/idea or if I had somehow accidentally manipulated it into assigning religious purpose to the task of unix-style streaming rpcs. I think a lot of the “clean” stuff stems from system prompts telling it to behave in a certain way or giving it requirements that it later responds to conversationally. Total aside: I actually really dislike that these products keep messing around with the system prompts so much, they clearly don’t even have a good way to tell how much it’s going to change or bias the results away from other things than whatever they’re explicitly trying to correct, and like why is the AI company vibe-prompting the behavior out when they can train it and actually run it against evals.
- Helmut10001 5mo agoI had the feeling they didn't really answer the questions, that is why the goblins appeared. They simply "retired the “Nerdy” personality" because they couldn't fix it and went on.
- isege 5mo agoOne I noticed with gemini, especially 3 flash: "this is the classic _____".
- afro88 5mo ago> The obsession with the word "seam" as it pertains to coding I quite liked this term when it started using it. And I appreciate the consistent way it talks about coding work even when working on radically different stacks and codebases
- creamyhorror 5mo ago"Seam" has been stretched by AI from its original legacy-code context to any point in code where something can be plugged in. I actually asked an AI about this a few weeks ago because I was surprised by the consistent, frequent use of "seam". Frequent words I see from GPT: "shape", "seam", "lane", "gate" (especially as verb), "clean", "honest", "land", "wire", "handoff", "surface" (noun), "(un)bounded", "semantics" (but this one is fair enough), and sometimes "unlock" It feels like AI really likes to pick the shortest ways to express ideas even if they aren't the most common, which I suppose would make sense if that's actually what's happening.
- joegibbs 5mo agoChatGPT has a whole host of weird words that it uses about coding - anything changed is a “pass” done over the code, it loves talking about “chrome” in the UI, it’s always saying “I’m going to do X, not [something stupid that nobody would ever think of doing]”
- bwat49 5mo agogpt also loves talking about handwaving, "I'm going to do X, not just a hand-wavy victory lap"
- duped 5mo agoShort terse sentences. Never use commas. Paragraph break. No foo. No bar. Only baz and qux. All writing is like a bad tech blog -- with language that mimics humanity. Yet is alien. The smoking gun is extra wording. Typically simple language. Dense in tokens -- shallow in content. Repeating itself ad nauseam. Saying the same thing in different ways. Feeding back upon itself. Not adding content. Not adding depth. Only adding words.