4 ms·
This article is complete bunk. The researchers used a "chatgpt detector" which as we've seen over and over in academia, do not work. This study is completely un
by chickenpotpie 3y ago
This article is complete bunk. The researchers used a "chatgpt detector" which as we've seen over and over in academia, do not work. This study is completely unfounded.
God I'm choking on the irony of an article about the dangers of using AI to train AI based on a study that used AI to detect AI
- foxbyte 3y agoI can understand your frustration with the article, but let's approach it with an open mind. While the use of a "chatgpt detector" may have its limitations, it's essential to appreciate the researchers' effort in exploring new methods. The study may not be perfect, but it contributes to the ongoing conversation about the risks of using AI in AI training. Irony aside, let's keep the discussion going and encourage further research to improve our understanding of this complex field.
- AndriyKunitsyn 3y agoSo polite it hurts. I wonder if in the future people on the internet will leave deliberately offensive posts to show that they are human.
- TeMPOraL 3y agoIt's crazy, isn't it? I don't know what feels worse for me - that whenever I read a mannered, well-structured and somewhat verbose comment, I now suspect it wasn't authored by a human - or that, as I quickly realized, my own writing style feels eerily similar to ChatGPT output.
- salt4034 3y agoIf it helps, your response here doesn't feel similar to ChatGPT output.
- TeMPOraL 3y agoThanks. I've already noticed that I've started to unconsciously adjust my writing style to avoid that feeling of similarity to ChatGPT. That said, compared to typical comments on-line (even on this site), using paragraphs, proper capitalization, correct punctuation, and avoiding typos already gets you more than half of the way to writing like ChatGPT...
- mvdtnz 3y ago@dang are we ever going to do anything about this? You almost can't read a comment section in a thread about AI without this crap now.
- parl_match 3y ago[flagged]
- sbierwagen 3y agoI mean, what rule do you actually want here? ChatGPT has been RLHFed into a pretty distinctive style, but there's no reason to think a better LLM wouldn't have a more natural style. If AGI is possible, then HN will end up with AI users who contribute on an equal basis to the modal HN user, and then shortly after that, more equal. Should all AI be banned? Should you have to present a birth certificate to create an account?
- parl_match 3y ago> Should you have to present a birth certificate to create an account? I actually honestly believe that the era of "open registration" forums and discussion places is going to come to a close, largely due to GNN. It's not going to become a problem until the hardware and walltime costs of training models and running them comes down. You'll know it's a problem when every 10th post on 4chan is a model pretending to be a human that is of a gentle but unyielding political persuasion of some sort. I don't know what the end pattern will be, but it'll likely be a combination of things - large platforms, like reddit or facebook, where individual communities "vibe check" posts out. or - some sort of barrier to entry, such as a small amount of money (the so called "idiot tax": if you're an idiot, you get banned, and you have to pay again) - some sort of (manual!) positive reputation system for discussion boards, sort of like how peering works - some sort of federation technology where you apply and subscribe to federation networks I don't think we'll really be able to predict what the future looks like right now (it's not even widely recognized as a problem). And since this is HN, I'll add: I don't think there's any serious money to be made running reputation or IDV, unless you've already started. And if it becomes a serious enough problem, players like ID.me/equifax/bureau will be the situation for "serious" networks (linkedin, facebook, chat, etc).
- jameshart 3y agoEveryone’s trying to take the shortcut. Can someone in this space invest in doing the hard work to have experts manually curate data? You know back before Wikipedia, publishers used to pay people to write and edit encyclopedias? It doesn’t scale. Sure. That’s what the AI you’re building is for though - it will scale. Throwing compute at ‘the entirety of the internet’ feels like such a lazy way to get what we’re after here.
- haney 3y agoThe company that I work at does exactly the service that you're describing. We recently spun up a team of Math PhDs to help with data labeling. (https://www.invisible.co/ https://www.invisible.co/). We're seeing more and more of our clients ask for graduate level data labelers and content creators.
- dontupvoteme 3y agoRight now I'm pretty sure just having gpt rewrite the average low to mediocre content that made up the gruel of its generic internet diet and doing fine tunes will get us another 5-10x along, but most hopefully for us little guys out there with a ~24-48GB VRAM cap If GPT4 really is 8 230M models, the next bit for us will be a few ~1-5M models that swap in for whatever you want to create, or talk about, or what have you Imagine a model trained just on English football for the purpose of having a good time in the pub that is used when the topic changes to it. I bet you could pass on the dailymails sports page if you add some "u"s into your words. Or a model finetuned specifically on the library you're trying to debug, maybe even specifically in combination with other tools you're trying to put together.
- benreesman 3y agoMy well-documented melancholy around the state of the LLM “conversation” notwithstanding, I’ll point out that there’s a long and generally productive history of adversarial training: from the earliest mugshot GANs to AlphaZero, getting these things to play against each other seems to produce interesting results. Whatever the merits of this or that “ChatGPT detector”, the concept isn’t unprecedented or ridiculous.
- lisasays 3y agoPer the article, the didn't just use the static detector: They also extracted the workers’ keystrokes in a bid to work out whether they’d copied and pasted their answers, an indicator that they’d generated their responses elsewhere. So while I don't yet know if the article is bunk -- I do know that your hot take is bunk.
- vminvsky 3y agoThey never used a static detector.
- lisasays 3y ago"Static" in this case refers its being used in isolation (not to any particular kind of detector). I could have said "they didn't just use the detector all by itself", I suppose.
- chickenpotpie 3y agoOr they typed up their responses in a different text editor and copy and pasted from that?
- TeMPOraL 3y agoPlot twist: the author will turn out to have enlisted the help of AI in writing this article.
- vminvsky 3y agoIt's actually pretty easy to create a bespoke ChatGPT detector!
- the_other 3y agoBut will it give reliable results?
- vminvsky 3y agoaccording to the paper they get 98% accuracy. another recent paper came out saying it's always possible to discriminate between real and synthetic text [1]. i think the core problem is with the generalist classifiers (gptzero, openai detector, etc). ex. openai's classifier has an accuracy of around 25% on it's own text. however, when you train a bespoke classifier (like the authors did), you can get really good results. [1] https://arxiv.org/pdf/2304.04736.pdf https://arxiv.org/pdf/2304.04736.pdf
- iinnPP 3y agoThe moment a detector is taken seriously is the moment it will be trivially beaten by another AI designed to beat the detector.
- vminvsky 3y agoi would recommend u read the paper. the contribution isnt a detector thats meant to be taken seriously; but a detector that works in a very specific task. they then use this to estimate use of LLMs on MTurk
- lisasays 3y agoIs it now? Adversarial training isn't infinitely scalable either, has its limitations also. Also - the moment that companies start training models to resist detectors, they expose themselves to regulation. Won't stop dark AI models running on some website somewhere, but it can be very effectively applied to companies running at Google or OpenAI scale.