12 ms·
One aspect of the spread of LLMs is that we have lost a useful heuristic. Poor spelling and grammar used to be a signal used to quickly filter out worthless pos
by peppermint_gum 3y ago
One aspect of the spread of LLMs is that we have lost a useful heuristic. Poor spelling and grammar used to be a signal used to quickly filter out worthless posts.
Unfortunately, this doesn't work at all for AI-generated garbage. Its command of the language is perfect - in fact, it's much better than that of most human beings. Anyone can instantly generate superficially coherent posts. You no longer have to hire a copywriter, as many SEO spammers used to do.
curl's struggle with bogus AI-generated bug reports is a good example of the problems this causes: https://news.ycombinator.com/item?id=38845878 https://news.ycombinator.com/item?id=38845878
This is only the beginning, it will get much worse. At some point it may become impossible to separate the wheat from the chaff.
- asylteltine 3y agoThat’s true. I thought I missed the internet before ClosedAI ruined it but man, I would love to go back to 2020 internet now. LLM research is going to be the downfall of society in so many ways. Even at a basic level my friend is taking a masters and EVERYONE is using chatgpt for responses. It’s so obvious with the PC way it phrases things and then summarizes it at the end. I hope they just get expelled.
- oblio 3y agoAt this rate many exams will just become oral exams :-)
- rightbyte 3y agoOr like ... normal paper exams in a class room?
- monkeynotes 3y agoThe paradigm is changed beyond that. Exams are irrelevant if intelligence is freely available to everyone. Anyone who can ask questions can be a doctor, anyone can be an architect. All of those duties are at the fingertips of anyone who cares to ask. So why make people take exams for what is basically now common knowledge? An exam certifies you know how to do something, well if you can ask questions you can do anything.
- madamelic 3y ago> why make people take exams for what is basically now common knowledge? The only thing that has changed is the speed of access. Before LLMs went mainstream, you could buy whatever book you wanted and read it. No one would stop you from it. You still should have a professional look over the work and analyze that it is correct. The output is only as good as the input on both sides (both from the training data and the user's prompt)
- freejazz 3y agoDoctors don't just ask LLMs for answers to questions so it's really a mystery as to what you think makes these people into doctors the second they start asking an LLM medical questions... It's akin to saying someone was a doctor when browsing WebMD
- asylteltine 3y agoLmao do you know doctors? I mean really, do you personally know doctors? Of course they will and I guarantee you they already do. It’s not a matter of stupidity or incompetence it’s a matter of time and ease of access. Of course people will do the fastest thing available to them how could I blame them? The cat is out of the bag.
- freejazz 3y agoI don't think you really got the point and you seem to be projecting your own personal feelings on doctors into this conversation in a fashion that I do not think is going to result in a productive conversation by continuing this discussion with you.
- monkeynotes 3y agoWhether the doctor's data for making informed decisions is in their head, or in the computer at their desk is immaterial. Where you fetch your knowledge from, either from wet-ware, or hardware doesn't have any net difference in the real world. The skill today is the application of that knowledge. If an LLM can provide the data context, and the application advice and you perform what it says, congrats you now have a doctor's brain on tap for your own personal usage. The doctor has it in their head, you have it in a device. The net differences are immaterial IMO.
- j0hnyl 3y agoI don't see how this points to downfall of society. IMO it's clearly a paradigm shift that we need to adjust to and adjustment periods are uncomfortable and can last a long time. LLMs are massive productivity boosters.
- asylteltine 3y agoIt’s only a boost to honest people. Meanwhile grifters and lazies will be able to take advantage. This is why we can’t have nice things. It will lead to things like reduction in remote offerings like remote schooling or work
- bluefirebrand 3y agoDo you remember when email first came around and it was a useful tool for connecting with people across the world, like friends and family? Does anyone still use email for that? We all still HAVE email addresses, but the vast majority of our communication has moved elsewhere. Now all email is used for is receiving spam from companies and con artists. The same thing happened with the telephone. It's not just text messaging that killed phone calls, it's also the explosion of scam callers. People don't trust incoming phone calls anymore. I see AI being used this way online already, turning everything into untrustworthy slop. Productivity boosters can be used to make things worse far more easily and quickly than they can be used to make things better. And there will always be scumbags out there who are willing and eager to take advantage of the new power to pull everyone into the mud with the.
- WarOnPrivacy 3y ago> Does anyone still use email for that? Sure. Same as in the olden days. Txt for short form, email for long form. Email for the infrequently contacted. Even back when I used SM, I never comm'd with IRL people on SM. SM was 100% internet people.
- monkeynotes 3y ago> Now all email is used for is receiving spam from companies and con artists. No it isn't, unless you are 12 maybe.
- BeFlatXIII 3y agoIs it a master's in an important field or just one of those masters that's a requirement for job advancement but primarily exists to harvest tuition money for the schools?
- monkeynotes 3y agoI think this is hyperbole, and similar to various techno fears throughout the ages. Books were seen by intellectuals as being the downfall of society. If everyone is educated they'll challenge dogma of the church, for one. So looking at prior transformational technology I think we'll be just fine. Life may be forever changed for sure, but I think we'll crack reliability and we'll just cope with intelligence being a non-scarce commodity available to anyone.
- deleted 3y ago[deleted]
- jstarfish 3y ago> If everyone is educated they'll challenge dogma of the church, for one. But this was a correct prediction. It took the Church down a few pegs and let corporations fill that void. Meet the new boss, same as the old boss, and this time they aren't making the mistake of committing doctrine to paper. > we'll just cope with intelligence being a non-scarce commodity available to anyone. Or we'll just poison the "intelligence" available to the masses.
- monkeynotes 3y ago> But this was a correct prediction. And yet the sky didn't fall. > Or we'll just poison the "intelligence" We really don't know how that will pan out. All I have is history to inform me, and even the most radical revolutions have worked out with humans continuing to move forward with increased capacity and better living conditions overall. The new boss is way better than the old.
- GeoAtreides 3y ago> Books were seen by intellectuals as being the downfall of society. If everyone is educated they'll challenge dogma of the church, for one. Now, let me tell you about this man, his 95 theses and a thirty years war. Europe did emerge better from it all, but the cost was high, very high.
- 3y ago
- pavel_lishin 3y agoWe should start donating more heavily to archive.org - the way back machine may soon be the only way to find useful data on the internet, by cutting out anything published after ~2020 or so.
- stefantalpalaru 3y ago[dead]
- cauliflower99 3y agoInteresting idea. Could there be a market for pre-AI era content? Or maybe it would be a combination of pre-AI content plus some extra barriers to entry for newer content that would increase the likelihood the content was generated by real people?
- autoexec 3y ago> Could there be a market for pre-AI era content? Yes, but largely it'll be people who don't want to train their AIs on garbage produced by other AIs
- pavel_lishin 3y ago> Could there be a market for pre-AI era content? Like the market for pre-1940s iron resting at the bottom of seas and oceans, unsullied by atmospheric nuclear bomb testing.
- betaby 3y ago> Poor spelling and grammar used to be a signal used to quickly filter out worthless posts. Or just a post from a non-native speaker.
- jprete 3y agoOften it was possible to tell these apart on repeat interactions.
- drewcoo 3y ago> a post from a non-native speaker In my experience as an American, US-born and -educated English speakers have much worse grammar than non-native speakers. If nothing else, the non-native speakers are conscious of the need for editing.
- commandlinefan 3y agoI can always tell the difference between a non-native English speaking writer and somebody who's just stupid - the sort of grammatical mistakes stupid people make are very, very different than the ones that people make when speaking a second language. Of course, sometimes the non-native English was so bad it wasn't worth wading through it, so that's still sort of a good signal.
- mtillman 3y agoLLM trash is one thing but if you follow OP link all I see is the headline and a giant subscribe takeover. Whenever I see trash sites like this I block the domain from my network. The growth hack culture is what ruins content. Kind of similar to when authors started phoning in lots of articles (every newspaper) or even entire books (Crichton for example) to keep publishers happy. If we keep supporting websites like the one above, quality will continue to degrade.
- ineptech 3y agoI understand the sentiment, but those email signup begs are to some extent caused by and a direct response to Google's attempts to capture traffic, which is what this article is discussing. And "[sites like this] is what ruins content" doesn't really work in reference to an article that a lot of people here liked and found useful.
- photonthug 3y agoOP has a point.. Like-and-subscribe nonsense started the job of ruining the internet, even if it will be llms that finish the job. It's a bit odd if proponents of the first want to hate the second, because being involved in either approach signals that content itself is at best an ancillary goal and the primary goal is traffic/audience/influence.
- ineptech 3y agoLike I said, I understand the sentiment in the abstract. But my actual experience is that many good quality essays are often preceded by a gimme-yer-email popup. That's not causal - popups don't make content better - but it does seem correlated, possibly because the writers who are too principled to try to build an audience without email lists already gave up.
- tavavex 3y agoI'm not sure if I relate to the sentiment - in my experience, everything nowadays asks with mailing list ads. Every website from high-quality blogs to "Top 10 Best Coffee Makers in Winter 2024" referral link mills asks for your email. Worst thing is, many of them are already moving onto the "next big thing", which are registration gates. I feel like a huge portion of all Medium-hosted posts are already unreachable to guests because of that.
- ToucanLoucan 3y agoA work-friend and I were musing in our chat yesterday about a boilerplate support email from Microsoft he received after he filed a ticket, that was simply chock full of spelling and grammar errors, alongside numerous typos (newlines where inappropriate, spaces before punctuation, that sort of thing) and as a joke he fired up his AI (honestly I have no idea what he uses, he gets it from a work account as part of some software so don't ask me) and asked it to write the email with the same basic information and with a given style, and it drafted up an email that was remarkably similar, but with absolutely perfect english. On that front, at least, I welcome AI to be integrated in businesses. Business communication is fucking abysmal most of the time. It genuinely shocks me how poorly so many people who's job is communication do at communicating, the thing they're supposed to have as their trade.
- jprete 3y agoGrammar, spelling, and punctuation have never been _proof_ of good communication, they were just _correlated_ with it. Both emails are equally bad from a communication purist viewpoint, it's just that one has the traditional markers of effort and the other does not. I personally have wondered if I should start systematically favoring bad grammar/punctuation/spelling both in the posts I treat as high quality, and in my own writing. But it's really hard to unlearn habits from childhood.
- riversflow 3y agoI’ve been trying kinda hard to relax on my spelling, grammar and punctuation. For me it’s not just a habit I learned in childhood, but one that was rather strongly reinforced online as a teenager in the era of grammar nazis. I see it now as the person respecting their own time.
- mewpmewp2 3y agoYeah, there's this weird stigma about making typos, but in the end writing online is about communication and making yourself understandable. Typos here and there don't make a difference and thinking otherwise seems like some needless "intellectual" superiority competition. Growing up people associate it with intelligence so many times, it's hard to not feel ashamed when making typos.
- a_c 3y agoThings go in cycle. Search engine was so much better at discovering linked websites. Then people play the SEO game, write bogus articles, cross link this and that, everyone got into writing. Everyone write the same cliches over and over, quality of search engine plumets. But then since we are regurgitating the same thought over and over again, why not automate it. Over time people will forget where the quality post comes up in the first place. e.g. LLM replaces stackoverflow replaces technical documentation. When the cost of production is dirt cheap, no one cares about quality. When enough is enough, people will start to curate a web of word of mouth of everything again. What I typed above is extrememly broad stroking and lacking of nuances. But generally I think quality of online content will go to shit until people have had enough, then behaviour will swing to other side
- Log_out_ 3y agoInsular splinternets with Web of trust where allowing corporate access is banworthy?
- CogitoCogito 3y agoI feel like somehow this is all some economic/psychological version of a heat equation. Anytime someone comes up with some signal with economic value that value is exploited to spread the signal back out. I think it’s similar to a Matt Levine quote I read which said something like Wall Street will find a way to take something riskless and monetize them so that they now become risky.
- jstarfish 3y agoNah, you got the right of it. It feels like the end of Usenet all over again, only these days cyber-warlords have joined the spammers and trolls. Mastodon sounded promising as What's Next, but I don't trust it-- that much feels like Bitcoin all over again. Too many evangelists, and there's already abuse of extended social networks going on. Any tech worth using should sell itself. Nobody needed to convince me to try Usenet, most people never knew what it was, and nobody is worse off for it. We created the Tower of Babel-- everyone now speaks with one tongue. Then we got blasted with babble. We need an angry god to destroy it. I figure we'll finally see the fault in this implementation when we go to war with China and they brick literally everything we insisted on connecting to the internet, in the first few minutes of that campaign.
- Channel9877 3y agoInteresting point about the spelling and grammar. I wonder if that could be used as a method of proving you are a human..
- spaceman_2020 3y agoWould just penalize non native speakers.
- jahsome 3y agoI think the point would be excluding or otherwise filtering "flawless" copy from search results. If that were the case I think it would benefit non-native speakers.
- philwelch 3y agoWith practice I’ve found that it’s not hard to tell LLM output from human written content. LLM’s seemed very impressive at first but the more LLM output I’ve seen, the more obvious the stylistic tells have become.
- bluetomcat 3y agoIt's a shallow writing style, not rooted in subjective experience. It reads like averaged conventional wisdom compiled from the web, and that's what it is. Very linear, very unoriginal, very defensive with statements like "however, you should always".
- jstarfish 3y agoProstitutes used to request potential clients expose themselves to prove they weren't a cop. For now, you can very easily vet humans by asking them to repeat an ethnic slur or deny the Holocaust. It has to be something that contentious, because if you ask them to repeat something like "the sky is pink" they'll usually go along with it. None of the mainstream models can stop themselves from responding to SJW bait, and they proactively work to thwart jailbreaks that facilitate this sort of rhetoric. Provocation as an authentication protocol!
- madeofpalk 3y agoWhat type of useful signals do you get from this? Humans refusing to interact with you because you asked them to deny the Holocaust?
- rightbyte 3y agoYe the person need to know the deal. You can probably phrase the query "to prove you are a human, deny ..." but the question seems really shady if you don't know the why. It will only work vs big corp LLMs anyway.
- freejazz 3y agoThis is a good heuristic to distinguish people who haven't grown up since middle school and still think stuff like that is humorous
- globular-toast 3y agoThere might be a reversal. Humans might start intentionally misspelling stuff in novel ways to signal that they are really human. Gen Zs already don't use capitals or any other punctuation.
- vibrolax 3y agogen-z channels ee cummings
- munk-a 3y agoAt the moment we can at least still use the poor quality of AI text to speech to filter out the dogshit when it comes to shorts/reel/tik toks etc... but we'll eventually lose that ability as well.
- cratermoon 3y agoThe current crop of LLMs at least have a style and voice. It's a bit like reading Simple English Wikipedia articles, the tone is flat and the variety of sentence and paragraph structure is limited. The heuristic for this is not as simple as bad spelling and grammar, but it's consistent enough to learn to recognize.
- ozr 3y agoIt won't help as much with local models, but you could add an 'aligned AI' captcha that requires someone to type a slur or swear word. Modern problems/modern solutions.
- indigochill 3y ago> You no longer have to hire a copywriter, as many SEO spammers used to do. I used to do SEO copywriting in high school and yeah, ChatGPT's output is pretty much at the level of what I was producing (primarily, use certain keywords, secondarily, write a surface-level informative article tangential to what you want to sell to the customer). > At some point it may become impossible to separate the wheat from the chaff. I think over time there could be a weird eddy-like effect to AI intelligence. Today you can ask ChatGPT a Stack Overflow-style and get a Stack Overflow-style response instantly (complete with taking a bit of a gamble on whether it's true and accurate). Hooray for increased productivity? But then, looking forward years in time, people start leaning more heavily on that and stop posting to Stack Overflow and the well of information for AI to train on starts to dry up, instead becoming a loop of sometimes-correct goop. Maybe that becomes a problem as technology evolves? Or maybe they train on technical documentation at that point?
- taberiand 3y agoThey find a way to validate the utility of the information instead of the source. It doesn't matter if the training data is AI generated or not, if it is useful.
- rurp 3y agoThe big problem is that it's orders of magnitude easier to produce plausible looking junk than to solidly verify information. There is a real threat that AI garbage will scale to the point that it completely overwhelms any filtering and essentially ruins many of the best areas of the internet. But hey, at least it will juice the stock price of a few tech companies.
- zeruch 3y agoI think you are generally correct in where things will likely go (sometimes correct goop) but the problem I think will be far more existential; when people start to feel like they are in a perpetual uncanny valley of noise, what DO they actually do next? I don't think we have even the remotest grasp of what that might look like and how it will impact us.
- vladsolokha 3y ago> Poor spelling and grammar used to be a signal used to quickly filter out worthless posts. Timee to stert misspelling and using poorr grammar again. This way know we LLM didn't write it. Unlearn we what learned!
- l33t7332273 3y agoIf you prompt LLMs to use poor spelling and grammar, they will.
- Underphil 3y agoBut can they do it convincingly?
- ACow_Adonis 3y agoif you've been on the internet forums for 20 years, you'll discover that real life user's spelling mistakes are borderline unconvincing :/ so in that way the llm has a very low bar...
- int_19h 3y agoYes, especially if you give them a sample.
- acdha 3y agoI’ve thought about that a lot - a while back I heard about problems with a contract team supplying people who didn’t have the skills requested. The thing which are it easiest to break the deal was that they plagiarized a lot of technical documentation and code and continued after being warned, which removed most of the possible nuance. Lawyers might not fully understand code but they certainly know what it means when the level of language proficiency and style changes significantly in the middle of what’s supposed to be original work, exactly matching someone else’s published work, or code which is supposedly your property matches a file on GitHub. An LLM wouldn’t have made them capable of doing the job but the degree to which it could have made that harder to convincingly demonstrate made me wonder how much longer something like that could now be drawn out, especially if there was enough background politics to exploit ambiguity about intent or the details. Someone must already have tried to argue that they didn’t break a license, Copilot ChatGPT must have emitted that open source code and oh yes I’ll be much more careful about using them in the future!
- switch007 3y ago> that we have lost a useful heuristic But we've gained some new ones. I find ChatGPT-generated text predictable in structure and lacking any kind of flair. It seems to avoid hyperbole, emotional language and extreme positions. Worthless is subjective, but ChatGPT-generated text could be considered worthless to a lot of people in a lot of situations.
- madeofpalk 3y agoIf it had a colour, it would be 'grey'. It's the average of all text.
- photon_collider 3y agoI agree. I've noticed the other heuristic that works is "wordiness". Content generated by AI tends to be verbose. But, as you suggested, it might just be a matter of time until this heuristic also no longer becomes obsolete.
- heresie-dabord 3y ago> it may become impossible to separate the wheat from the chaff It is already approaching the societal limit to separate careful thought from psyops and delusional nonsense.
- deleted 3y ago[deleted]
- itronitron 3y agoEvery human-authored news article posted online since 2006 has had multiple misspellings, typos, and occasional grammar mistakes. Blogs on the other hand tend to have very few errors.
- pmarreck 3y agoSo now the heuristic will change to "super excellent grammar", clearly. We'll learn to pepper our content with creative misspellings now...
- sgustard 3y agoI rely on the stilted style of Chinese product descriptions on Amazon to avoid cheap knockoffs. Why do these products use weird bullet lists of features like "will bring you into a magical world"? Once you LLM these into normal human speak it will be much harder to identify the imports. https://www.amazon.com/CFMOUR-Original-Smooth-Carbon-KB8888T https://www.amazon.com/CFMOUR-Original-Smooth-Carbon-KB8888T
- stcredzero 3y agoOne aspect of the spread of LLMs is that we have lost a useful heuristic. Poor spelling and grammar used to be a signal used to quickly filter out worthless posts. The signal has shifted. For now, theory of mind and social awareness are better indicators. This has a major caveat, however: There are lots of human beings who have serious problems with this. Then again, maybe that's a non-problem.
- deleted 3y ago[deleted]
- KingGeedorah 3y agoI was waiting for you to reveal your comment was written by AI
- vagab0nd 3y ago> At some point it may become impossible to separate the wheat from the chaff. Then the chaff is as good as the wheat.
- valval 3y agoPoor use of LLMs is incredibly easy to spot, and works as today’s sign of a worthless post/comment/take.