6 ms·
AI-generated content, other unfavorable practices get CNET on Wikipedia banlist
- kgeist 3y agoInteresting that the article itself gives AI-generated vibes: >It's important to remember that while Wikipedia is "The Free Encyclopedia that Anyone Can Edit,"
- EGreg 3y agoEnjoy for now that you can still tell. It’s early times like the video games before Call of Duty etc. The worst thing about AI is how it can easily betray you, manipulate you and have swarms execute long term sleeper plans at scale!
- batch12 3y agoSounds like all information we're given should be met with a healthy dose of skepticism. The ability to smell AI makes it a little easier right now to put it in the garbage pile.
- passion__desire 3y agoWe are moving toward a future where future robotics viruses will infect deployed humanoids and command them to go on a killing spree. https://www.youtube.com/watch?v=WWAnJX889j0 https://www.youtube.com/watch?v=WWAnJX889j0
- lolinder 3y agoEh, ChatGPT uses phrases like "it's important to remember" because they were common in the training data. The rest of the sentence definitely isn't ChatGPT: > It's important to remember that while Wikipedia is "The Free Encyclopedia that Anyone Can Edit," it's hardly The Wild West. My reading is that the article is average human prose (not great, not unreadable), not LLM prose.
- EGreg 3y agoIf I play GM-level chess moves from Rybka or AlphaZero randomly on average every 5 moves on average, could you tell?
- AndyNemmity 3y agoEvery 5 moves, yes. But as most Chess Grandmasters say, you only really need to use it in one difficult spot to change the result of a game.
- ummonk 3y agoNah. I play the top engine move like half the time and I'm not a particularly strong player.
- staunton 3y agoBoth can be true at the same time. A lot of the time the best move is obvious. However, if you play the top move 20% of the time, you will also play the top move in a lot of cases where it's not obvious. Given enough games, it's detectable.
- seanmcdirmid 3y agoI wonder if ChatGPT was trained on too many high school essays. It’s not very good writing style, but it’s what most people start out writing.
- jameshart 3y ago> My reading is that the article is average human prose. Which is, of course, exactly what ChatGPT is trained to produce. A lot of people's mental AI detector is actually a mediocrity detector.
- famouswaffles 3y agoNo GPT is trained to spit out anything that falls into its training data distribution. There's no average. Pre-training incentives being able to predict the smartest string of text in the corpus as readily as the dumbest. It doesn't converge on "average" and it doesn't really make sense that it would either. Base models don't talk like GPT. This is strictly an artifact of post training fine-tuning/RLHF.
- wdb 3y agoHow would you recognise AI-generated content? Is it a particular writing style? As a dyslexic person I am having a hard-time recognising it
- kgeist 3y agoChatGPT loves to add "it's important to remember" to every second answer. Also "multifaceted" and a few other adjectives like that.
- mistrial9 3y agoobjection -- ChatGPT does not "love" anything.. a reason to repeat that phrase is that it is written often, for real communication reasons. Declaring that something written often by humans is now an indication of ChatGPT is wrong-direction IMHO and directly objectionable here, a place where reasoning humans think and communicate via writing.
- kgeist 3y agoChatGPT was additionally finetuned after the initial training of the base model, so it's not completely clear if it truly represents the original distribution anymore. I use open weights models and they repeat such phrases less frequently in my experience.
- redox99 3y agoThat's completely untrue. Try using any base (non finetuned model) such as llama base, mistral base etc and you'll see that it does not write like that. That response format and writing style was added by OpenAI during the finetuning stage.
- mistrial9 3y agothe post says "a reason to repeat that phrase" not THE reason, not THE ONLY reason, so it is not COMPLETELY untrue. agitated?
- 3y ago
- shagie 3y ago(disclaimer: I'm a casual user of ChatGPT and haven't gone too much into the how it works)00 Given its nature as an LMM and a complex next word predictor, the phrase "it is important to remember" could be a way that was inadvertently trained so that it keeps itself on track. If it has "some points, it is important to remember {something}, some more things" it may be able to better generate text compared to "some points, weird tangent". Since it doesn't have a hidden memory, everything that it "thinks" is out there in the text including its own cues for what it should do. It also can't go back an edit its previous text to remove the self hints or clarify earlier points without calling them out. That style of writing differs from natural human writing since we are able to keep on topic (or not) without needing to write messages to ourselves that others can read. When we do, it's pointed out rather than trying to slip it in casually. "Note to self or reader" for things that are to be pointed out and break the flow of the text or "as an aside" for the tangents.
- coldblues 3y agoGiven that, would a chain of thought + final answer be a better output then?
- deleted 3y ago[deleted]
- ChrisArchitect 3y ago[dupe] Some more discussion here: https://news.ycombinator.com/item?id=39556041 https://news.ycombinator.com/item?id=39556041
- dang 3y agoSince that thread didn't get significant attention, the current thread doesn't count as a dupe. There are some comments there though: Wikipedia downgrades CNET's reliability rating after AI-generated articles - https://news.ycombinator.com/item?id=39556041 https://news.ycombinator.com/item?id=39556041 - Feb 2024 (10 comments)
- ChrisArchitect 3y ago2 days is not that long ago. 30+ upvotes and 10 comments is plenty of eyeballs for a recent topic. Conversely, it's not a dead topic there since it's only a few days, so the discussion could remain there? The point is there's already a recent post that many saw, but we have to see it again instead of just continuing the conversation there? <Shrug> Definitely a Related: of course.
- dang 3y ago30 upvotes and 10 comments is borderline, yes, but that thread only spent 8 minutes on HN's front page. If it had spent more time there I'd agree. And Btw you can check this kind of thing (more or less) using https://hnrankings.info/39556041/ https://hnrankings.info/39556041/ - it didn't catch the 8 minutes but the approximation is good enough. I suppose I could've merged the current comments thither and re-upped that one, but I didn't think of it!
- MyFirstSass 3y ago[flagged]
- mysterydip 3y agoMaybe a resurgence of local news outlets? "news written by people you can trust to be people" might be what we need to bring that subscription model back.
- atrus 3y agoI'm not sure what changes if this brilliant essay is written by bot, human, or martian?
- ianmcgowan 3y agoA connection to truth, or objective reality if you still believe in that.
- Zenst 3y agoA good read is a good read - whoever writes it. Whole issue of fake media and that, well, look at social media over the years, I doubt it will make much difference than what we have already. People over time, learn what to trust and who to trust. The real issue though, as you touch upon, will be governments using this as another way to gain even more power and control of our daily lives. Though my thoughts are, like most new tech that can copy/mimic - it will be the Disney's and movie studios with their lobbyist money that will do the most harm in any progress. Though mindful of the potential damage, a rouge AI could do in the financial markets. One upside - postal letters will start to see an increase in usage as people start to trust physical media more over digital. Is certainly a trend I foresee happening.
- brucethemoose2 3y agoAI is unique in that the incentive is to write engaging, compelling, perhaps wonderful ostensibly truthful content that is total nonsense. Enthusiasts in a subject are often wrong, but usually the goal is at least to convey something useful about the topic to the reader. The spammer/content farmer doesn't care about that. Users trying to farm attention on social media don't care about that. A manipulator on a certain topic doesn't care about that. They just want SEO and eyeballs. But usually the intersection between high quality topical writers and people with these motivations is very low. ...That is not the case any more.
- GuB-42 3y agoSee https://xkcd.com/978/ https://xkcd.com/978/ (Citogenesis) for what I think is the biggest problem with a source that publishes AI-generated articles. Wikipedia is supposed to use primary sources, AI generated articles, by nature, can't be primary sources. In particular, AIs love to use Wikipedia in their training dataset: it is a free, high quality source of information, but it is not flawless either. So there is a good chance that if Wikipedia cites an AI generated article, it has Wikipedia as its source, starting the "citogenesis" process.
- shh2112 3y agoSlight nitpick, while Wikipedia allows primary sources in some cases, it generally prefers secondary sources. https://en.wikipedia.org/wiki/Wikipedia:No_original_research#:~:text=A%20primary%20source%20may%20be,but%20without%20further%2C%20specialized%20knowledge https://en.wikipedia.org/wiki/Wikipedia:No_original_research....
- jMyles 3y agoThat's more than a slight nitpick - it corrects a core misunderstanding to be found everywhere this discussion seems to be happening. To wit: AI articles are to be found along a dramatic spectrum of quality. If such an article is high-quality, relevant to the subject matter, asserts the fact in question, and makes proper and veracious use of a primary source in support of such an assertion, why isn't it a reasonable source for an encyclopedia? I can imagine a future with rich educational materials with these layers: * raw experimental data -=> * publication ("primary" source) -=> * AI-generated review of many publications ("secondary" source) -=> * Human-authored encyclopedia article (with one or two people following the fact pattern all the way back down to the data, and many more people helping to synthesize the higher layers into a rich, readable, considerate, diverse synopsis)
- easton 3y agoI thought Wikipedia loved secondary sources? Like articles about something, as opposed to a primary source which would be like a person who was present or something (which would mean it’s original research). An article that says Abraham Lincoln wrote the emancipation proclamation would be preferred to Abraham Lincoln himself being asked and then the Wikipedia article citing “my interview with Abe”. (Your general point is valid, I’m just confused about the terminology)
- 2four2 3y agoIt'll be interesting if AI-generated content is ever no longer distinguishable from human writing in both evidence and prose. I suppose that's the ultimate goal. The final hurdle at that point will be the de-democratization of writing and the dilution of creativity and novel writing. There will probably always be a market for that, but for things like reporting events, it seems like AI could easily overtake the industry.
- userbinator 3y agoThe more important question is whether it's no longer distinguishable because the AI-generated content has improved, or because human-generated content has regressed.
- DanHulton 3y agoI'm not sure if you mean this for real, or if you're just doin' it for the joke, but if you _are_ being serious, perhaps you're just reading the wrong things? There are some truly excellent books, articles, blog posts, and more being written by actual humans out there today, probably more than ever before.
- ukuina 3y agoSeparating signal from noise is going to become much harder. Here's a spoiler-filled article from a major site (IGN) whose last paragraph (and likely more) has clearly been authored by ChatGPT in order to meet tight publishing deadlines and coincide with a movie's release: https://www.ign.com/articles/dune-part-2-post-credits-scene-ending-explained-sequel https://www.ign.com/articles/dune-part-2-post-credits-scene-... The indicators are all there: "In short, (vague sentence)... Is (hypothetical from article) or (other hypothetical from article)? Will (vacuous statement) or (inverse vacuous statement)? We'll have to wait for (eventuality) to find out." There is no attribution to AI for the generated content, though, and this lack of attribution is going to become the norm once LLMs become just another authoring tool like spell-check. Coupled with the race for clicks, the "excellent blog posts" are going to be drowned out.
- 3y ago
- deleted 3y ago[deleted]
- smittywerben 3y agoWait, you're telling me their malware-infested downloader isn't a reliable source for news, either?
- ametrau 3y ago> The site that was hurt by this so-called SEO heist is called Exceljet, a site run by Excel expert David Bruns to help others better use Excel. Welcome to the enshittening. There’s only so much a single dedicated operator can take before they pack it in. We need legislation to catch up fast and some big symbolic restitution cases decided in the courts.