5 ms·
I think this is a lazy criticism that I am _also_ growing tired of. If LLMs are trained on written information, that pattern of speech was present before they
by bcraven 11mo ago
I think this is a lazy criticism that I am _also_ growing tired of.
If LLMs are trained on written information, that pattern of speech was present before they got there. It's a good way to add emphasis.
- trueno 11mo agoI don't think anyone's here to debate the origin of speech patterns these things are using. Feels clear to me at least that the guy you're replying to is uninterested in reading stuff generated by AI, I can't say I disagree with him.
- stingraycharles 11mo agoI’m very much uninterested in reading AI generated content. Your assertions seems to be “AIs only write like that because people have been writing like that”, but that’s not a great argument. It feels like AI has suddenly given a platform for people who previously were unable to properly write blog content. But it immediately feels unoriginal and generic. I’m just not interested in that type of content and immediately put off by it. The only reason I mentioned this is because of the comment about Gemini 3 being in the comments. I’m just really, really tired of all the AI content everywhere nowadays and crave some authenticity. It just feels like cheap remakes / imitations to of original content.
- deleted 11mo ago[deleted]
- tankenmate 11mo agoFor the most part LLMs choose "the most common" tokens; so regardless of whether the content was "AI content" or not, maybe you are getting tired of mediocrity. And of course also that mediocrity has now become so cheap that it is now the overwhelming majority.
- stingraycharles 11mo agoLLMs have the tendency to really like comparisons / contrasts between things, which is likely due to the nature of neural networks (eg “Paris - France + Italy” = “Rome”). This is because when representing these concepts as embeddings, they can be computer very straightforward in vector space. So no, it’s not all due to human language, LLMs do really write content in a specific style. One recent study also showed something interesting: AIs aren’t very good at recognizing AI generated content either, which is likely related; they’re unaware of these patterns. https://www.sciencedirect.com/science/article/pii/S1477388025000131 https://www.sciencedirect.com/science/article/pii/S147738802...
- catlifeonmars 11mo agoThis is similar to how the average number of children per household is 2.5, but no one has 2.5 children. The most common tokens actually yield patterns that no one actually uses together in practice
- bmacho 11mo agoWas present, so what? It was 1 in 1 million, now it's 999999999 in 1 million. It is perfectly valid getting really tired of it, in fact, this is exactly what "getting really tired of" means and has always meant.
- NicuCalcea 11mo agoCertain patterns are much more common in LLM output than in human writing. I'm a journalist and love an em dash, for example, but I've never met/read another journalist that uses them nearly as often as LLMs. Same with the "this isn't just X, it's Y" pattern. When you have multiple of these patterns in every paragraph, it's a pretty clear indicator that the text is AI-generated. Plus, the author admitted to using AI to write it.
- Tanoc 11mo agoOne of the little tics I've noticed that helps weed out and LLM generated text is to CTRL+F and look for the word "therefor" in any of it. LLMs will use the word in a new sentence that isn't the conclusion of any previous sentences or paragraphs. Think like, "Bees are small fuzzy and yellow. Therefor their ability to fly is an astounding achievement." In all of my years of reading I've never seen people use the word that much in common writing, and when they do it's usually as part of a compound sentence. These things really do have their own little set of semantics and dialect that they follow that seems like it's a unique quirk.
- op00to 11mo agoNo, that’s in literally every LLM generated response to a forum message I’ve seen. It’s so common as to have become a trope. That’s not confidence. It’s a clear indicator of AI.
- krsdcbl 11mo agoI'll second that, this is extremely annoying and exhausting. It feels like the slightest occurrence of a less-than-ubiquitous pattern or any word not regularly used by the majority of the population instantly spawns a sleuth of newfound linguists who'll pitch in to explain how this certain marker ought to be proof of AI origin. This does nothing for the conservation, except helping the claim that AI will erode and dumb down our language become a self-fulfilling prophecy when people start feeling pressured to use the most dumbed down, simplistic and rhetorically bland way of expressing themselves to avoid any "suspicion"
- a2128 11mo agoNot necessarily, the LLMs used today are far from just simple models of written information on the internet. They use in-house data they wrote themselves, and RLHF/DPO where it's effectively training on its own data to optimize for human preference. If sampling with high enough temperature for this, it could theoretically bring out entirely new unseen forms of speech as long as people express their preference for it via the user interface