6 ms·
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual con
by mlsu 2mo ago
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter.
- "Introduction that rephrases your prompt."
- "3 paragraphs, with one section of bullet points"
- "The Twist"
- "The Bottom Line"
It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane observation about California burritos, is phrased in exactly the same way. This is obviously an artifact of post-training but it's also kind of how you can tell that this thing is a lot closer to a blindsight scrambler than real intelligence.
- dominotw 2mo agoanyone know how they would train a model to have this proclivity ?
- zucked 2mo agoThe second bullet point, down to the comma in the middle of the sentence, is what has been driving me absolutely batty of late. It's a surefire tell that I cannot seem to beat out of my outputs. It CONSTANTLY does it, even when you say not to. Between that and the insistence on "this, not that" structure makes me want to install the caveman skill and use it even for non-code workflows.
- craigmcnamara 2mo agoThis load bearing concern belt and braces.
- nonethewiser 2mo agoProvenance
- LVB 2mo agoSeam
- alexchantavy 2mo agoPressure test
- EdwardDiego 2mo agoMy personal bugbear is its usage of "grain" where normally you'd use "granularity", if at all.
- ltrg 2mo agoYes! Also everything's a "gate" that needs to be "wired up".
- paradox460 2mo agoLet me get my fence pliers
- smj-edison 2mo agoThat's the smoking gun
- EdwardDiego 2mo agoIs it the smoking gun? Or is it the twist? Or maybe, even, the trap?
- sevenseacat 2mo agoLately I've noticed everything is "sharp" or it "sharpens" the point or question or something.
- perryizgr8 2mo agoThat's a sharp question, and it reframes the whole point.
- dr_dshiv 2mo agoBelt and braces all the way down. They did their post training in British context?
- HappMacDonald 2mo agoI always get "Belt and suspenders" from Fable, so I can't tell
- phickey 2mo agoIn your reading, what is the distinction between blindsight’s scrambler and real intelligence? My reading is that it’s just as real, and draws out the disadvantages a sense of self constrains intelligence with
- pmontra 2mo agoI think that Blindsight's scramblers were intelligent but not conscious, which for us is very difficult to understand. For them consciousness was a blight. For the HN readers that are missing the context: https://www.rifters.com/real/Blindsight.htm https://www.rifters.com/real/Blindsight.htm Full text on the web site of the author.
- mlsu 2mo agoSure. I should have been more precise about what is 'real intelligence' here. What I mean is that blindsight's scramblers are aliens that cannot share human values. Their structure is completely different to ours, their qualia (or whether they even have it) is impossible for us to understand. In short, they do not have a soul. When Claude does this "slowly revealing a dramatic insight" thing that it does, it does that not because it has judged itself through some introspection as having an insight to share. It does not even know what an insight is or is not. It is not sharing anything, because it is not capable of sharing, because it does not have a soul. The aesthetic structure of its replies is a pattern, a constraint on the token distribution, like the color of noise. It's my bad to use the word 'intelligence' because it's so overloaded. Will Claude will act as a therapist or produce value or produce a work of art? No. It cannot, because it does not have a soul. I leave it freely open to interpretation whether having a soul is required for "real intelligence." But what I've noticed is that "intelligence" in these discussions is mostly used to denote some capability to produce [economic/social] value. In my mind value is a relational thing, a thing of human feeling.
- xyzzy_plugh 2mo agoI think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence." I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed. When we choose words it isn't because of a token distribution, nor because we rolled a die. We choose words because we feel a certain way, the external world, our body and senses are all connected as one system. These machines don't experience moods or get tired or feel better after a good night's sleep. They don't know their audience, we're all the same to them. We have no personal relationship nor can we establish one, as presenting some arbitrary background is not the same thing as a fluid, evolving relationship that accumulates through experience over time. There are no scars or fond memories. If these things can truly be intelligent, to abuse your use of the word, then at least we are quite far from holding them correctly.
- herbturbo 2mo agoPerhaps this is related to their new "invisible watermark" concept which would probably require rather contrived language patterns to make possible.
- nonethewiser 2mo agoI think opus was released before they included it on model. Its hard to say, but from what Ive read it doesn’t seem like it would have that drastic of an effect. I had the same thought though.
- rhdunn 2mo agoThe watermarking is independent of the model. The model itself has the probability weights to determine the next token. The watermark is similar to things like temperature and top_p/top_k in that the watermark adjusts the probabilities in a deterministic way that changes over time to hide tells from word choices.
- fcarraldo 2mo agoI love this theory. "We've invented a new invisible watermark that can detect whether code is LLM written." The watermark: counting instances of 'load-bearing seam', 'the hard truth', 'and that's the whole point'.
- cma 2mo agoIf it's using Aaronson's approach it shouldn't have any noticeable affect on generations. When it picks between options weighted by probability after the generation of logits, it still follows the probability mass, it just uses a known pseudorandom seed so that when you go back and look at the exact choices you can fingerprint it.
- FabHK 2mo agoThat's exactly right. And as is well understood, a good pseudorandom generator, despite being fully deterministic, is extremely hard to distinguish from randomness, unless you have the algorithm and key (internal state). Quite smart, really.
- beAbU 2mo agoIt's the same way how every AI generated poster looks exactly the same. As if there is a single underlying prompt that describes the template of the poster/long-form article, and it does not dare deviate from that.
- edoceo 2mo agoIsn't there? Like everyone using $MODEL is starting from the same base system-prompt. Then our user input is a small bit on top of that core mode. Like what would happen if everyone asked Mikey to paint their ceiling - they'd all be similar and therefore boring.
- kanzure 2mo agoAt this point, I'm basically telling models to not write any English text or prose. Only write code. They are great at writing code. Not so great at writing good English. In software projects, lengthy comments and docs are an anti-pattern: the software should instead be written to do the right expected thing so that you don't have to think about it. I don't want all these tokens polluting my context, either.
- alertchecker 2mo agoAgreed, think the first user rule I ever put into Cursor was "Don't write code comments unless absolutely necessary to explain something that couldn't just be inferred"
- boomlinde 2mo agoI've noticed that ChatGPT (whatever model the free version uses by default) likes to phrase answers as though it's correcting me, even when my question doesn't contain any assumptions.
- bborud 2mo agoDoesn’t just mean the model has picked up on what happens when two graybeards who still wear cargo shorts meet and one utters a declarative sentence? :-)
- buu700 2mo agoInterestingly, I've been noticing almost the opposite issue. 5.6 Sol frequently starts its responses with "Yes" even when my prompt doesn't contain a yes-or-no question.
- timpera 2mo agoI've been noticing the same thing since 5.4- it starts with "Yes" yet I haven't asked a question. I think it might be related to the reasoning, like it's answering its own questions?
- ssl-3 2mo agoIt's been happening for years, and it does seem to be entirely related to its background reasoning leaking out of the context and into the output. The "It's not X, it's Y"-style repetition is another example of that: It argues with itself in the background, and the argument leaks out as if in refutation of things that you (the human) never had imparted into discussion at all.
- smoe 2mo agoSomething I've noticed quite a bit on the paid plans as well is that it starts its answers with "I mostly agree..." or "almost correct...," then goes through the list of points I made without actually disagreeing with any of them. I assumed this is a system prompt or RL that nudges it to be always skeptical but then it still has like all the models the urge to appease the user.
- Panoramix 2mo agoYou're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
- mlsu 2mo agoYou are diabolical.
- domoregood 2mo agoWell played.
- IgorPartola 2mo agoI am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us. And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated. And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.
- onlyrealcuzzo 2mo ago> I am curious why LLM writing has such an uncanny valley feel to it. Because they are HEAVILY trained to give addictive responses. They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.
- 2mo ago
- visarga 2mo agoI think the Opus 5 formula is to be the little professor treating your ideas like an essay for grading, or like a buyer analyzing merchandise for purchase.
- paradox460 2mo agoYou're not just right, you're correct!
- beezlewax 2mo agoAnd this is infuriating. I don't want to read all this gibberish anymore. It's making me hate what software engineering has become.
- xivzgrev 2mo agoYes! Ive started getting a feel for AI writing on blogs. It feels slightly verbose and involves "reveals" "It's not the naked man on your lawn waving a chainsaw that's scaring you. It's the burrito you ate for lunch: it went down easy, but now it's coming for you"
- userbinator 2mo agoProbably trained on lots of clickbait.
- chr15m 2mo ago> The aesthetic is that of an expert slowly revealing an insight to the user. Ah, that's it! Thank you. I wonder if they are training it to talk like this because that's what their customers actually want? They want a machine genius to lead them.
- nlawalker 2mo agoIt’s the TED Talk playbook: the crafting of a lecture given by an expert to laypeople to maximize attention, engagement and satisfaction. Every piece of prose is built to pack in as many TED Talk mic-drops/expectation-subverting insight bombs as possible.
- lukebouch 1mo agoI think you nailed it with your explanation! That is exactly it.
- Wowfunhappy 2mo ago...I mean, on the whole, I'm glad it's detectable. I imagine they could have post-trained it to not be detectable.
- felipeerias 2mo agoThe model is generating tokens one by one and that sentence structure allows it to keep its options open rather than committing at the beginning of the sentence
- ldng 2mo agoWhat an amazing way to inflate token spending ! /s
- itemize123 2mo agoexactly, I found that it's eerily similar to those marketing copies. gets obnoxious
- torarnv 2mo agoThis twist often involves loosely related or even unrelated bugs or even non-issues, and Claude bringing those into the conversation at that point breaks my mental processing of the response. Any tips on how to avoid that would be highly appreciated!
- theptip 2mo agoA related theory here is that Opus is heavily RL’d to be a sub-agent. If Fable is the primary interlocutor then perhaps there is less pushback on the obtuse language. Indeed perhaps the convoluted language acts as a kind of Neuralese between models deriving from the same pretrained base.
- nfw2 2mo agoIt's an artifact of the reasoning process I think. Setting reasoning effort to none works better when trying to change writing style ime
- miguelspizza 2mo agoThe more I use the AIs the more I feel like I did at the end of reading blindsight. The things are undoubtedly “intelligent” by any practical definition of the word, but they are not aware. This is also why I advocate against using AI as a writing partner. No matter the argument you lay out, a fresh context window will always have the “a few good things and a few bad things” feedback. There is no higher order opinion to align with.
- lenglain 2mo agoAm I crazy for thinking that this is a pretty big regression compared to past models? I remember being blown away by GPT 4.5, and I kept using it up until they decommisioned it. I think claude 3.7 sonnet was pretty good too. Gemini seems to be the best one right now for actually talking. Opus is top tier for code but when i talk to it I want to rip my hair out. GPT-5.6 is doing best for me right now among the powerful models.