12 ms·
GPTZero Case Study – Exploring False Positives
- Maken 4y agoAlternative title: Most academic papers are indistinguishable from AI generated babble.
- sigmoid10 4y ago...to AI. It's kinda funny how this is yet another area where these models suck very much in the same way that most humans do. LLMs are bad at arithmetic? So are most people. Can't tell science from babble? I already wouldn't ask a non-expert to rate any aspect of an academic paper. Trusting the average Joe who has only completed some basic form of education would be tremendously stupid. Same with these models. Maybe we can get more out of it in specific areas with fine tuning, but we're very far away from a universal expert system.
- brookst 4y agoThe best was the Ted Chiang article making numerous category errors and forest/trees mistakes in arguing that LLMs just store lossy copies of their training data. It was well-written, plausible, and so very incorrect.
- dr_dshiv 4y agoI felt the same way. But I’d love to read a specific critique. Have you seen one?
- elefanten 4y agoHere’s one from a researcher (which also links to another), though I’m not qualified to assess it’s content in depth. https://twitter.com/raphaelmilliere/status/1624073150475431940 https://twitter.com/raphaelmilliere/status/16240731504754319...
- supriyo-biswas 4y agoNeural network based compression algorithms[1] are a thing, so I believe Ted Chiang's assessment is right. Memorization (albeit lossy) is also how the human brain works and develops reasoning[2]. [1] https://bellard.org/nncp/ https://bellard.org/nncp/ [2] https://www.pearlleff.com/in-praise-of-memorization https://www.pearlleff.com/in-praise-of-memorization
- brookst 4y agoThe fact that some neural network architectures can compress data does not mean that data compression is the only thing any neural network can do. It’s like saying that GPUs can render games, so GPT is a game because it uses GPU.
- jejeyyy77 4y agoI mean, humans can't distinguish AI written text either - which is why this tool was built? I don't see how it will be possible to build such a tool either as the combination of words that can come after another is finite.
- Maken 4y agoI do agree that the most likely reason is that scientific papers tend to be highly formulaic and follow strict structures, so a LLM is be able to generate something much more alike to human writing than if it tries to generate narrative. But it's still fun to deduce that the reason is that the quality of technical writing has sunk so low, that is even below the standards for AI generated text.
- jldugger 4y ago>...to AI. Perhaps we can call it the "Synthromorphic principle," the bias of AI agents to project AI traits onto conversants that are not in fact AI.
- deleted 4y ago[deleted]
- nobu-mori 4y agoThat's great! All we need to do is negate the output and it will be more accurate.
- explaininjs 4y agoNon-editorialized title: GPTZero Case Study (Exploring False Positives) @dang
- deleted 4y ago[deleted]
- icapybara 4y agoRight, I mean, how could it know? I could write in the tone of ChatGPT if I tried hard enough. It's an intractable problem and a tool like this probably does more harm than good.
- JoeAltmaier 4y agoIf it's like many of these 'AI' engines, it's a statistical map of text. I'd expect it to produce no better than what it's trained on, then muddy that by randomly combining different versions of similar statements. I'm impressed it is ever accurate.
- brookst 4y agoI’ve got a fun little side project that uses GPT. I tested gptzero against 10 of my projects’ writings and 10 of my own. It detected 6 out of 10 correctly in both cases (4 gpt-written bits were declared human, 4 human-written were declared gpt). Which is better than 50% but not nearly good enough to base any kind of decision on.
- d0mine 4y ago> which is better than 50% Unrelated: p-value for getting 12 from 20 correct just by chance is ~0.4 that is there is not enough data for the conclusion "better" in this case. Null hypothesis: 50%/50%, the result random, normal distribution: H0: p=1/2 H1: p!=1/2 (two-tail) import statistics p0 = 0.5 # proportion of successes according to null hypothesis n = 20 # sample size p_sample = 12/n # 12 from 20 are correct sigma = (p0 * (1 - p0) / n)**.5 # std according to H0 z_score = (p_sample - p0) / sigma # test statistic p_value = 2*statistics.NormalDist().cdf(-abs(z_score)) # prob. two-tails # p-value -> 0.4
- panarky 4y agoAs millions of people interact with ChatGPT, their writing will subtly, gradually, begin to mimic its style. As future versions of the model are trained on this new text, both human and AI styles will converge until any difference between the two are infinitesimal.
- onos 4y agoInteresting idea, but isn’t there variance in the output? Eg I’ve seen people ask it to “write in the style of x” etc and different people also clearly have different writing styles.
- Filligree 4y agoYou can try, but it isn’t very good at that. The style remains very ChatGPT.
- moralestapia 4y ago+1 We train AIs but they also train us.
- ominous 4y agoSome related idea, in case you like to see that thought explore: https://medium.com/@freddavis/we-shape-our-tools-and-thereafter-our-tools-shape-us-1a564cb87484 https://medium.com/@freddavis/we-shape-our-tools-and-thereaf...
- moralestapia 4y agoGreat read, thanks for sharing!
- brookst 4y agoOne of the big complaints with LLMs is the confident hallucination of incorrect facts, like software APIs that don’t exist. But the way I see it, if ChatGPT thinks the Python list object should have a .is_sorted() property, that’s a pretty good indication that maybe it should. I work in PM (giant company, not Python), and one of these days my self-control will fail me and I will open a bug for “product does not support full API as specified by ChatGPT”.
- wand3r 4y agoI saw this[1] interview with Sam Altman touching on interim AI impact. I really agree with his point that basically detecting output from LLMs is basically going to be futile and only really relevant in the near term. Accuracy is obviously going to improve in models and detection isnt that difficult now but will be in the future, especially if output is modified or an attempt to obfuscate origin is made. [1]https://youtu.be/ebjkD1Om4uw https://youtu.be/ebjkD1Om4uw
- sebzim4500 4y ago>detection isnt that difficult now I would have thought this, but every attempt I've seen at detecting chatGPT generated text has failed miserably.
- Closi 4y agoIt fails on false positives, but you don't often get false negatives, which might be enough at the moment for quite a few use-cases. Also false positives are typically "this text is likely to contain parts that were AI generated" rather than "This text is higly likely to be AI generated" (which is what GPT-generated content generally produces). When I've tried to prompt-engineer GPT to produce text that GPTZero will flag as negative it has been pretty tough!
- theptip 4y agoOf course false positives matter; the naive heuristic “everything is AI generated” has zero false negatives, and mostly false positives. In the OP half of the positives are false. That’s not a useful signal IMO. You couldn’t use that to police homework for example.
- Closi 4y agoI said they might not matter for quite a few use cases, not that they don't matter for all use-cases. e.g. If you were Sam Altman at OpenAI and your use-case is mostly looking for training data and wanting to tell if it is AI-Generated or not (so you can exclude this from training data), you probably care much more about false negatives than false positives (false positives just reduce your training data set size slightly, while false negatives pollute it). Of course they matter if you are marking homework (where conversely false negatives aren't actually that important!), but it's pretty trivial to think of use-cases where the opposite is true.
- wongarsu 4y agoThis is testing on abstracts of medical papers. I wouldn't be surprised if the way abstracts are carefully written and reviewed by > 3 people is very different from normal text, and in some ways more similar to LLM output.
- jmfldn 4y agoIt's important to shine a light on the limitations of AI detection software, and this case study on GPTZero does just that. False positives can have serious consequences, particularly in sensitive areas such as healthcare.
- MonkeyMalarky 4y agoThank you for your heroic effort in copying and pasting a chat log, I am left in awe by it. ...nice edit...
- jmfldn 4y agoLighten up mate.
- bentcorner 4y agoRandom thought: In the future the simplest method to reduce your chance of GPT usage being detected is to pretend your paper was written in 2022 or earlier.
- k__ 4y agoYes, it's really bad. False positives and negatives all over the place. I really wish it worked, but generally, it doesn't.
- lwhi 4y agoPerhaps AI generated text should be created with a specific signature in mind _specifically_ to be identifiable?
- mattnewton 4y agoisn’t this essentially asking anyone who runs a model to flip the evil bit[0]? People who want to misrepresent model output as human written output will trivially be able to beat this protection by removing the signature or using a version of the model that simply doesn’t add it. [0] https://en.m.wikipedia.org/wiki/Evil_bit https://en.m.wikipedia.org/wiki/Evil_bit
- lwhi 4y agoI think you're correct, but it would promote the idea that what's produced is a basis or source for further work .. and would mean that effort is required. Feels like a good basis tech for something like ChatGPT.
- wongarsu 4y agoThere's a large body of research into invisible text watermarking, so this would certainly be possible. Maybe the simplest to implement in LLMs would be to bias the token generation slightly, for example by making tokens that include the letter i slightly more likely. In a long enough text you could then see the deviation from normal human text characteristics.
- wankle 4y ago"Tell me more about your perspective customers." https://projects.csail.mit.edu/films/aifilms/AIFilms.html https://projects.csail.mit.edu/films/aifilms/AIFilms.html
- netfortius 4y agoFoxNews is accurate 10% of the time, at best, and it's allegedly produced by humans, so I'm not really seeing the problem here...
- ClumsyPilot 4y agoyou do it all wrong, it's 90% accurate and 90% sucessfull - its goal to be as inaccurate as possible without its audience realising. If fox news deletws the last 10% of reality from it's broadcasting, the people might start catching on. Although these days i am not sure
- dsign 4y agoI had some fun yesterday when ChatGPT hallucinated a bibliographic reference to an article that didn't exist. But the journal existed, and it had plenty of articles that made ChatGPT's hallucination plausible. I think that at least this use case can be fixed with some pragmatic engineering[^1]. [^1]: Which may take a bit to happen, because our current crop of AI researchers have all taken "The bitter lesson"[^2] to heart. [^2]: http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- FPGAhacker 4y agoDid chatgpt post the reference as a footnote (or parenthetical)? At least for now,I was thinking it didn’t do that, and maybe the lack of references would be an indicator of unedited gpt output.
- dsign 4y agoIt did: ... > However, according to a study published in the Journal of Dairy Science, the diacetyl content of butter can range from approximately 0.5 to 14 parts per million (ppm) (https://doi.org/10.3168/jds.S0022-0302(03)73775-3 https://doi.org/10.3168/jds.S0022-0302(03)73775-3). The doi I could not find, so I'm pretty sure is bogus. I asked it to produce a full reference: > Sure, the full reference for the study I mentioned is: > Yvon, M., Chambellon, E., & Bolotin, A. (2003). Effect of pH on diacetyl and acetoin production by Lactococcus lactis subsp. lactis biovar diacetylactis. Journal of dairy science, 86(12), 4068-4076. I went to the index of the journal 86(12), and that article is not there.
- int_19h 4y agoIt's okay, GPT just gets confused about which of the quantum many-worlds it is currently in. Just ask it to write the article as needed. ~
- ozy23378 4y ago[dead]
- photochemsyn 4y agoJust wrote this myself, although I did try to chatGPT-style it a bit. I thought the final third would serve to identify it as non-AI as it goes off on a tangent about isotopes... > "The periodic table is a systematic ordering of elements by certain charcteristics including: the number of protons they contain, the number of electrons they usually have in their outer shells, and the nature of their partially-filled outermost orbitals." > "Historically, there have been several different organizational approaches to classifying and grouping the elements, but the modern version originates with Dmitri Mendeleeve, a Russian chemist working in the mid-19th century." > "However, the periodic table is also somewhat incomplete as it does not immediately reveal the distribution of isotopic variants of the individual elements, although that may be more of an issue for physicists than it is for chemists." GPTZero says: "Your text is likely to be written entirely by AI" Now I'm feeling existential dread... perhaps I am an AI running in a simulation and I just don't know it?
- thomastjeffery 4y agoTechnical writing, in order to be relatively unambiguous - the "technical" part - defines itself as a subset of English with a constrained grammar and vocabulary. You just illustrated technical writing. Naturally, your writing style is very similar to that of other technical writing. Take one guess what kind of writing exists in most of the text GPT is trained on.
- catchnear4321 4y agoHow has your “success” been with chatgpt? Qualitatively, generally, positive/negative. Being able to speak to machines would likely correlate with “sounding like one.” There could be a different (mis)categorization here but also por que no los dos.
- LoveMortuus 4y agoI mean, technically speaking you are a man made intelligence, thus it would be fair to say that you are artificial intelligence. ^^
- sdwr 4y agoYou made a few unforced errors that move it away from the quietly authoritative, mirror-sheen AI voice. The colon in line 1 is clunky, the combination of "but" and "with" in line 2 reads as passive, and line 3 is full person.
- mt_ 4y agoIsn't the GPTZero based on model detector by OpenAI for GPT-2? In the initial preview it was exaclty like the demo, a text box styled diferently. [1] - https://github.com/openai/gpt-2-output-dataset/blob/master/detector/README.md https://github.com/openai/gpt-2-output-dataset/blob/master/d...
- gtsnexp 4y agoDetecting text generated by large language models like ChatGPT is a challenging task. One of the main difficulties is that the generated text can be highly variable and can cover a wide range of topics and styles. These models have learned to mimic human writing patterns and can produce text that is grammatically correct, semantically coherent, and even persuasive, making it difficult for humans to distinguish between the text generated by machines and the ones written by humans. Another challenge is that large language models are highly complex and constantly evolving. GPT-3, for example, was trained on a massive dataset of text and can generate text in over 40 languages. With this level of complexity, it can be challenging to develop detection systems that can keep up with the ever-changing text generated by these models. To implement a reliable detection system like GPTZero, which is designed to detect text generated by GPT-3, several challenges need to be addressed. First, the system needs to be highly accurate and efficient in identifying text generated by GPT-3. This requires a deep understanding of the underlying language model and the ability to analyze the text at a granular level. Second, the system needs to be scalable to handle the vast amounts of data generated by GPT-3. The detection system should be able to analyze a large volume of text in real-time to identify any instances of generated text. Finally, the system needs to be adaptable to the evolving nature of large language models. As these models continue to improve and evolve, the detection system needs to keep up and adapt to the changing landscape.
- ericmcer 4y agoIt was weird how easy this was to identify if you have read any amount of ChatGPT content. It has a particular writing style that is pretty obvious. I am not sure how you would code something to detect an author based on writing style. It feels like something people would have tried to do before. Probably using a similar approach that LLMs use but with a separate predictor for specific authors.
- deleted 4y ago[deleted]
- mds 4y agoChatGPT in particular writes in middle school essay format: introduction, point 1, point 2, point n, conclusion.
- corrupto 4y ago[flagged]
- janalsncm 4y agoGPTZero cannot work because they don’t have all of the logits used to judge perplexity. They only have the top five or so. Even OpenAI which has the full set of logits is not able to correctly classify all texts. It’s a futile effort.
- jaimex2 4y agoWho didnt see this coming when GPTZero was announced by clueless media? I didn't even have to look at the internals to know GPTZero and OpenAI's solutions would be pointless. I'm still very surprised OpenAI made snake oil.
- air7 4y agoI read an interesting paper about an idea of watermarking LLM output text in such a way that makes detection very accurate for a long enough text. This is done by subtly changing the probabilities of the next word to be generated based on the last word that was outputed. Circumventing it by manually changing words post hoc would potentially require almost as much work as writing it from scratch. The idea seems quite roboust to me and I can envisage a future where companies that provide access to LLMs would also publish a detection tool for their models.
- tluyben2 4y agoIt's not hard to make a model that rewrites the text without changing the meaning which fails this. Our model[0] which is based on feeding chatgpt random things from the interwebs from before 2020 and letting it wobble on about it is pretty good and nice to play with, but it's pretty easy to change the score radically with just changing a few words. This is whack-a-mole no matter how it's done. For now you can be sure people are too lazy to do it, but there will be many tools in the future to evade tools like this. [0] https://filteroutai.com/validate/a07e081b71b294ba2de236441beef118302fd3e895d4b82b490d697fd62e3888 https://filteroutai.com/validate/a07e081b71b294ba2de236441be... https://filteroutai.com/validate/2c3fa6de32845df02be7a4ff185a6c2c99c93e4b056ba49fb0b5d3322ea5c224 https://filteroutai.com/validate/2c3fa6de32845df02be7a4ff185...