21 ms·
33-46% of workers on MTurk used LLMs in a text production task
- paxys 3y agoI wonder how many of those tasks were for training LLMs.
- CoastalCoder 3y agoOut of curiosity, what happens to an LLM if, over time, it's increasingly trained on its own output and/or that of other LLMs? And has anyone figured out a good way to minimize that when sourcing training data from e.g. Reddit?
- hospadar 3y agoI certainly have no idea but interesting to think about! One could argue that humans are trained on their own output so presumably interesting things could arise, especially if it's not a totally closed loop (LLM#1 -> some human curation -> LLM#2).
- fladrif 3y ago> humans are trained on their own output And it seems we're emulating that same cycle. Humans are trained not only on our own output, but also take changes from our environment. So in this sense LLM's environment would be human input.
- cyanydeez 3y agothe problem with the current tech is that the solid part of knowledge isn't separate from the fluid parts. You need to fundamentally understand that human knowledge in many generations was purposely fixed, by books, scholars, etc. Until that occurs, it's really just a miasma of bullshit.
- pjmorris 3y agoI read (probably on HN) where someone suggested that pre-LLM data will become the 'low background steel' (pre-Nuclear-age steel with lower levels of internal radiation) of machine learning. It seems like provenance/traceability of data sources will become important, both for the builders and users of LLMs.
- verytrivial 3y agoThere's a no-longer-very-speculative bit of Sci-Fi waiting here about whole societies "forking" based upon the curation of their training data. Arguably social-media + bots + paid state actors means we're already there.
- mrbungie 3y agoIsn't "societies forking based upon curation of their training data" practically what culture is? Now, I agree that having scalable generative AI and high-throughput-high-virality mass media (social networks) brings new unspeakeable horrors like feedback loops to the mix. Interesting times indeed.
- cyanydeez 3y agocan you point to an actual paid state actor. the term seems like mostly bullshit.
- deleted 3y ago[deleted]
- xwdv 3y agoContext drift.
- jmount 3y agohttps://news.ycombinator.com/item?id=36319076 https://news.ycombinator.com/item?id=36319076
- spacemanspiff01 3y agoI believe that performance degrades, as the models end up training on data that is more homogenized and less diverse when compared to original data. At least according to this paper (I think, FYI I am not a expert) https://arxiv.org/abs/2305.17493 https://arxiv.org/abs/2305.17493
- jerpint 3y agoIt can reinforce its own biases for one
- MengerSponge 3y agoHave you seen Multiplicity? "Hey Steve, come on up, I'm spittin on bugs"
- gnicholas 3y agoHm, sounds like dog-fooding, but where the 'food' is dog shit. Dog-shitting? Dog-shit-fooding?
- gpderetta 3y agoDog breakfasting
- lordnacho 3y agodogging
- TheSpiceIsLife 3y agoDogenshittification
- jstarfish 3y agoThe Poopboros.
- Daishiman 3y agoAI Centipede
- asow92 3y agoHopefully some of this is democratized away by humans voting on quality of output directly and/or indirectly?
- lozaning 3y agoHabsburg AI – a system that is so heavily trained on the outputs of other generative AI's that it becomes an inbred mutant, likely with exaggerated, grotesque features from https://twitter.com/jathansadowski/status/1625245803211272194 https://twitter.com/jathansadowski/status/162524580321127219...
- samstave 3y agoOr even, the Zuckerberg of AI in a meta-sorta-way'' This was pretty clever: >CaligulAI <-- @Rogntuudju
- CoastalCoder 3y ago> CaligulAI I believe it's spelled "Caligulae" because it's first declension. /Latin-joke
- samstave 3y agoLOVE it You know what would be funny ; ."Latin Joke Explainer" -- As explained in Latin Man-Splain-Terms."* EDIT: (You mastered my joke, and I appreciate it.) EDIT AGAIN: UI think we actually just coined a term ; 'Caligul::AI' -- Malicious AI for its own pleasure.
- Mistletoe 3y agoYou’ve heard it at a concert when the microphone picks up sound from the speakers and a positive feedback loop ensues. A lot of screeching.
- veave 3y agoLately I have been trying a new LLM called Falcon that was supposedly created in the UAE but it was mostly trained on conversations with ChatGPT... the result is that if you ask Falcon who created it, it will happily say OpenAI. I found that really amusing.
- whimsicalism 3y agoThe instruct trained variation, you mean.
- Aerbil313 3y agoI read a research paper which says LLMs don’t actually improve by training on the outputs of other LLMs. They only gain the ability to answer the specific questions included in the LLM output training dataset. Their reasoning capabilities etc. don’t improve.
- Aerbil313 3y agoI’m probably wrong. See the Orca model which came out recently.
- cj 3y agoPerhaps the source training data will shift to transcripts or subtitles of podcasts, cable television, TV shows, and other sources that haven't (yet) been polluted. Combined with pre-LLM data sets.
- avereveard 3y agoPossibly such data will be used in a different way. If it's embedded in a social network, engagement (votes, retweets,wwillhstecer) will act as a sort of crowd sourced reinforcement learning source instead of being direct part of the training set.
- Balooga 3y ago"Model Collapse" https://venturebeat.com/ai/the-ai-feedback-loop-researchers-warn-of-model-collapse-as-ai-trains-on-ai-generated-content/ https://venturebeat.com/ai/the-ai-feedback-loop-researchers-...
- the8472 3y agoTraining is becoming multi-modal anyway and will need grounding in physical reality so human-generated text will most likely become a smaller fraction of the data.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- jmount 3y agoThat is hilarious, and it utter defeats one of the quality checks on MTurk style tasks: using agreement as a proxy for correctness.
- Choco31415 3y agoYou could potentially introduce an occasional prompt and check the answer against a few LLMs to see if they match.
- adventured 3y agoGood LLM outputs are typically not the same for every identical query. For now that makes AI checkers fairly incompetent (and leads to disastrous teacher results when trying to find the students using AI). If I ask even just GPT 3.5 something like: "Who was John Adams?" - it'll give me a slightly varied answer pretty much every single time (even if I modify its settings to make it less creative with responses, it'll still usually vary a bit). Here is a simple API hit on gpt-3.5-turbo-16k-0613 Output 1) John Adams was an American statesman, lawyer, diplomat, and Founding Father who served as the second President of the United States from 1797 to 1801. He was one of the key figures in the American Revolution and played a crucial role in drafting the Declaration of Independence. Adams also served as the first Vice President under George Washington. He was known for his strong advocacy of republicanism and his belief in a strong central government. Adams was a prolific writer and his letters and writings provide valuable insights into the early years of the United States. Output 2) John Adams was an American statesman, lawyer, diplomat, and Founding Father who served as the second President of the United States from 1797 to 1801. He was one of the key figures in the American Revolution and played a crucial role in drafting the Declaration of Independence. Adams was also a strong advocate for the separation of powers and a strong central government. He was known for his intellect and commitment to public service. And so on.
- samstave 3y agoWhat if every output from a prompt was a QR code which led the user to a page with the actual prompt and the output and this is what AI was relegated to do in order to ad some modicum of proof-of-prompt-custodianship... so the only way one could see the output of the AI was to read the QR code which resulted in the intput & the output?
- ghaff 3y agoNot really a surprising result. It's a task tailor-made for LLMs and probably somewhat painstaking for a human--especially one that isn't a domain expert which the typical MTurk presumably wouldn't be.
- jerpint 3y agoIn hindsight it’s not surprising but when you’re paying turkers you expect humans to be doing the work, good to know it’s no longer reliable either
- ghaff 3y agoTotally. It's a useful result if only to confirm that people doing tasks that can be handled pretty well by automation will use said automation. Not surprising, but absent a test it's just a supposition based on human behavior. While less problematic, I expect that human transcribers are probably also using ML systems to generate a first pass (though I'm not sure how much that benefits a fast touch typist).
- AndrewKemendo 3y agoThe Turing Syndrome speeds up! Turing Syndrome (Like the Kessler syndrome) is when the amount of AI generated data on the internet surpasses human generated content to the extent that it will eventually make it impossible to distinguish between the two
- adventured 3y agoThat's not how it's actually going to play out. As the AI generated content becomes dramatically overwhelming in scale, the human content will become increasingly easy to spot (and there will be multiple cues to the human content that make it fairly obvious). There will be a crossing of the two along the way, in regards to the amount of content generated, a relatively brief time where it will be difficult to tell which is which. The more AI content there is, and the more it advances, the easier it's going to be to play spot the human. The time in which it'll be hardest to tell them apart, will be in the middle frames rather than in the later stages.
- 317070 3y agoBecause the human data will become "stupid" in comparison, is that the reason? I don't understand why spot-the-human will become easier otherwise.
- HideousKojima 3y agoThe human content will contain the n-word or other shibboleths.
- chaxor 3y agoThis also doesn't make any sense, as it would be much cheaper to use local LLMs to generate text, making it likely to see generated toxic language.
- jerf 3y agoThe "n-word" or any arbitrary other shibboleths may guarantee you didn't use OpenAI services, but they don't have any sort of lock on LLM technology and it's trivial to ask non-controlled ones to use any you like.
- PeterStuer 3y agoThe onlybthing that sprang to mind: 'Why so low?'
- none_to_remain 3y agoHow long till OpenAI et al. offer a plagiarism detection service against this?
- giobox 3y agoOpenAI are already attempting, they have a free classifier you can try, but results aren't all that great: "Our classifier is not fully reliable. In our evaluations on a “challenge set” of English texts, our classifier correctly identifies 26% of AI-written text (true positives) as “likely AI-written,” while incorrectly labeling human-written text as AI-written 9% of the time (false positives). Our classifier’s reliability typically improves as the length of the input text increases." https://openai.com/blog/new-ai-classifier-for-indicating-ai-written-text https://openai.com/blog/new-ai-classifier-for-indicating-ai-...
- none_to_remain 3y agoBut they don't need a fancy AI classifier, they already have the text that ChatGPT outputs
- thrtythreeforty 3y agoThey'd be absolute fools not to be keeping a copy of everything it's ever said. How much text has it generated, ever? A couple of dozen terabytes? Maybe?
- johntb86 3y agoAccording to https://help.openai.com/en/articles/7730893-data-controls-faq https://help.openai.com/en/articles/7730893-data-controls-fa... if you turn off history the conversation will be permanently deleted after 30 days.
- ghaff 3y agoAssuming you're actually going to do anything with the results, basically any meaningful number of false positives is a huge problem. If using an LLM (or appearing to have done so) to at least help write a college essay is considered plagiarism, that can easily be grounds for expulsion.
- dxbydt 3y ago33-46% of students in the local high school used LLMs in a text production task. I mean, they actually did. [1] https://www.fastcompany.com/90841387/gpt-3-chatgpt-high-school-schoolwork https://www.fastcompany.com/90841387/gpt-3-chatgpt-high-scho... [2] https://www.theatlantic.com/technology/archive/2022/12/openai-chatgpt-writing-high-school-english-essay/672412/ https://www.theatlantic.com/technology/archive/2022/12/opena... [3] https://www.nytimes.com/2023/01/12/technology/chatgpt-schools-teachers.html https://www.nytimes.com/2023/01/12/technology/chatgpt-school...
- bequanna 3y agoI am generally not very supportive of restricting AI research via regulation. However, I would support a mandatory disclaimer indicating that content was created using LLM. Consider it a “truth in sourcing” rule.
- Aerroon 3y agoA blanket regulation like this would be difficult to manage, because you're dealing with text. It might very well be generated in another jurisdiction that does not have such a rule. Even the service presenting you the text might not be aware that it's generated.
- bequanna 3y agoSeems like it could be enforced like GPDR? I actually don’t have a good idea on how this would be practically enforced. But I think we could start somewhere.
- ftxbro 3y agothey use mturk in LLM and other generative AI research as operationalized human evaluation
- bravura 3y agoSo mechanical turk, which was a robot that actually contained a human, is actually a robot that contains a human that contains a robot. We have achieved complete ouroboros.
- bhuga 3y agoOr perhaps "ourorobos" :)
- cscurmudgeon 3y agoAnd the final robot has millions-billions of humans fused using linear algebra.
- Terr_ 3y agoAnd the humans are trillions of tiny nanobots (after a there was a "Grey Goo" event billions of years ago) that formed a hivemind that isn't even aware of all its own parts. :P
- cscurmudgeon 3y agohttps://users.ece.cmu.edu/~gamvrosi/thelastq.html https://users.ece.cmu.edu/~gamvrosi/thelastq.html
- treprinum 3y agoIt's robots all the way down!
- nashashmi 3y agoWe call that reinforcement. It’s like actors in movies copy real world people who copy actors in movies.
- jeron 3y agoRLHBARF - Reinforcement Learning from Human But Actually Robot Feedback
- remote_phone 3y agoWhat will happen is that it will stratify people even more and income inequality will become even more pronounced. The educated people will learn how to write on their own and the less educated will use LLMs and never rise above it. That will be the McDonald’s workers of the Information Age. This is the future that I’m preparing my kids for so that they land on top. I’m investing heavily in creativity and writing for them so that they will land in the upper levels of society and not the lower ones that become slaves to AI.
- B1FF_PSUVM 3y agoBullshit generation as a service. Well, at least this makes it stand out how much of that was going on manually before it was automated.
- brd 3y agoMTurk has had this issue for years. If you tried to us MTurk to do any text or image labeling since at least 2018 you were majority of the time getting output from a poorly functioning ML system.
- SpaceManNabs 3y agowe are doomed to the fake web theory.
- antaviana 3y agoFunny because I remember in an AWS conference that the presenter touted Mechanical Turk as AAI (Artificial Artificial Intelligence) because it was a fake AI service done by humans, and it seems that it will become soon AAAI (Artificial Artificial Artificial Intelligence, or A3I).
- Imnimo 3y agoWhile the general conclusion of this paper seems perfectly plausible, I'm not sure how much weight I can put into the exact numbers they come up with. They only gather 46 summaries from MTurks, and they train a predictor on some known-good and known-synthetic data, and use that to decide when the crowd workers have submitted LLM-generated text. I'm not totally sold on the idea that the good performance on the validation set is going to translate to this new dataset. In general, I'm just really skeptical of results that say "we can detect generated text with 99% accuracy!". It would be very interesting to see a study that asks crowd workers afterwards if they used an LLM to complete the task. Obviously you'd be pretty worried about lying, but maybe there's a way to structure the situation so that the workers know their response is anonymous and won't affect whether they get paid.
- WesternWind 3y agoA friend was joking that LLMs explain star trek's obsession with the 20th and 21st century.
- SomeBoolshit 3y agoSo when these models start learning from more and more AI generated content, can we assume they'll get worse again and eventually make themselves useless?
- cyanydeez 3y agoI'd argue they've never been "good". They're at the level of random blogs on myspace.
- SnorkelTan 3y agoI had a side project I worked on for 6 months several years ago to do something similar albeit slightly less shady. When I was working on it there was a business on their that was using it to do human transcription of retail store receipts. I think it was a marketing company to figure out what customers were purchasing. My plan was to develop a computer vision algorithm that would do 90% of the work and then just farm out the verification step to humans to increase their productivity. I got as far as downloading the receipts and attempting basic image processing to try and touch up/reformat the receipts. Training a model to extract data would have been a better approach but I didn't have the skillset or inclination to continue.
- SnorkelTan 3y agoAlso, the revenue was there for some nice side income (~40k/mo), but not enough to support a company around it. I could have probably used services from an AI/ML startup doing this but it would have eaten into the margins.
- Animats 3y agoThere was an article yesterday indicating that feeding the output of LLMs back into them as training data causes the LLM to erase the original content. That effect, combined with people using Mechanical Turk to do pre-processing of content for LLM training, will make things worse. I can see repositories of pre-2022 textual content becoming valuable. Anything later will have too much circular corruption to be used as training data.
- mike_hearn 3y agoIt's this sort of thing that makes me bullish on OpenAI, oddly enough. I keep hearing rumours that it's possible to watermark LLM output with a sort of unique key such that the resulting text still sounds just as fluent, but it can be detected with 100% accuracy given more than a few words and the key. I also hear that OpenAI have mastered this technology. If that's true then it leads to the possibility of converting first mover advantage to absolute market dominance: 1. Spend lots of money to give a high quality LLM away for free (ChatGPT). Rapidly achieve market dominance by doing so. 2. Watermark its output without making a detector publicly available. 3. Users flood the internet with watermarked text. 4. Now you can generate fresh training data without circular corruption because almost all the LLM output on the internet is yours. 5. Your competitors incorporate LLM output into their training inputs because they can't detect it. Their quality suffers in ways yours don't. 6. Users notice the quality gap and continue to use your model. Competitors can't close it because they have no way to do so. 7. Profit! This sort of loop happened with Google where using the click stream to do ranking yielded a virtuous cycle that competitors struggled to beat, and then sites using robots.txt to ban non-Googlebots made it worse. If this equivalent actually happens then OpenAI can end up having an absolute monopoly on LLM technology.
- kmeisthax 3y agoThis is plausible, but I can put some bounds on the scope of the conspiracy. OpenAI has already released their own AI text classifiers[0], and they are nowhere near good enough to suggest the use of stylometric[1] watermarking. Furthermore, if such a watermark did exist, it would be detectable by anyone with a large enough corpus of known AI text. It probably would also be easier to detect (have lower false classification rates) than the unintentional stylo in ChatGPT output that most AI plagiarism detection is currently identifying. I could see OpenAI archiving literally every line of text that their models generated, and then doing normal full-text search on their text corpus to remove their own output from it. But there's still one other flaw: the existence of other models. GPT-3+ are not a monopoly. Google has BERT, Facebook has LLaMA, and the FOSS community has BLOOM, StableLM, and CerebrasGPT. While these models may or may not be GPT-4 quality, they will still be generating text without OpenAI's watermark, meaning that OpenAI won't know to filter for it. So even if they did plan to use watermarking as a way to invisibly scrape human content while polluting other companies' training sets, they wouldn't be able to succeed on such a plan. [0] https://openai.com/blog/new-ai-classifier-for-indicating-ai-written-text https://openai.com/blog/new-ai-classifier-for-indicating-ai-... [1] Stylometry is the process of fingerprinting word usage to determine authorship. This was used, for example, to unmask "Robert Galbraith" as the pen name of J.K. Rowling, before she made it too easy by just using her mystery novels as a way to rant about trans people and Twitter. More famously, the Unabomber was unmasked as Ted Kaczynski using the same means. Anonymity is a polite fiction and how you write is personally identifying information.