12 ms·
I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to dete
by capableweb 3y ago
I'm glad that they did, although they should obviously done an announcement for it.
The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could guarantee that, and act on whatever these "AI vs Human-written" tool tell them to.
So hopefully this can serve as another example that it's simply not possible to detect if a bunch of characters were outputted by an LLM or not.
- catboybotnet 3y agoThere's also the post going around about how it can (and does) falsely flag human posts as AI output, particularly among some autistic people. About as useful as a polygraph, no?
- capableweb 3y agoBoth false-positives are as useful as the other one, flagged "human" but actually "LLM" vs flagged "LLM" but actually "human". As long as no one put too much weight on the result, no harm would have been done, in either case. But clearly, people can't stay away from jumping to conclusions based on what a simple-but-incorrect tool says.
- frumper 3y agoA tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. There is no way to know when it’s wrong so you’ll either use it to help justify what you want, or ignore it. Edit: this tool is as reliable as a magic 8-ball
- dontreact 3y agoIf you were trying to predict the direction a stock will move (up or down) and it was right 99.9% of the time, would you use it or not?
- a13o 3y agoThis is a strawman. First, the AI detection algorithms can't offer anything close to 99.9%. Second, your scenario doesn't analyze another human and issue judgement, as the AI detection algorithms do. When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.
- dontreact 3y agoIt's not a strawman. There are many fundamentally unpredictable things where we can't make the benchmark be 100% accuracy. To make it more concrete on work I am very familiar with: breast cancer screening. If you had a model that outperformed human radiologists at predicting whether there is pathology confirmed cancer within 1 year, but the accuracy was not 100%, would you want to use that model or not?
- frumper 3y agoIt's a strawman because they aren't comparable to AI detection tests. A screening coming back as possible cancer will lead to follow up tests to confirm, or rule out. An AI detection test coming back as positive can't be refuted or further tested with any level of accuracy. It's a completely unverifiable test with a low accuracy.
- dontreact 3y agoYou are moving the goalposts here. The original claim I am responding to is "A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process." I agree that there are places where we shouldn't put AI and that checking whether something is an LLM or not is one of them. However I think the sentence above takes it way too far and breast cancer screening is a pretty clear example of somewhere we should accept AI even if it can sometimes make mistakes.
- frumper 3y ago
- bgirard 3y ago> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. It can be used for some decision (i.e. not critical ones), but it should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard. > this tool is as reliable as a magic 8-ball Citation needed
- frumper 3y agoThe AI tool doesn't give accurate results. You don't know when it's not accurate. There is no accurate way to check its results. Who should use a tool to help them make a decision when you don't know when the tool will be wrong and it has a low rate of accuracy? It's in the article.
- bgirard 3y ago> The AI tool doesn't give accurate results. Nearly everything doesn't give 100% accurate results. Even CPUs have had bugs their calculation. You have to use a suitable tool for a suitable job with the correct context while understanding it's limitation to apply it correctly. Now that is proper engineering. You're partially correctly but you're overstating: > A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. That's totally wrong and an overstated position. A better position is that some tools have such a low accuracy rate that they shouldn't be used for their intended purpose. Now that position I agree with it. I accept that CPUs may give incorrect results due to a cosmic ray event, but I wouldn't accept a CPU that gives the wrong result for 1/100 instructions.
- hn_go_brrrrr 3y agoThis is an unreasonable standard. Outside of trivial situations, there are no infallible tools.
- frumper 3y agoYou're right. After reading what I'd wrote, there should be some reasonable expectations about a tool, such as how accurate it is, or what are the consequences to be wrong. The AI detection tool fails both as it has a low accuracy and could ruin someones reputation and livelihood. If a tool like this helped you pick out what color socks you're wearing, then it's just as good as asking a magic 8-ball if you should wear the green socks.
- arcticbull 3y agoSeems a tautology no? “As long as we ignore the results the results don’t matter.”
- ImprobableTruth 3y agoFlagged "human" but actually "LLM" is not a false positive, but a false negative.
- WillPostForFood 3y agoIt depends how the question is framed: are you asking to confirm humanity, or confirm LLM. If you are asking, is this LLM text Human generated, and it says Human (yes), then it is false positive. If you are asking is this LLM generated text LLM generated, and is says and it says Human (no), then it is a false negative.
- NoZebra120vClip 3y agoThat seems like a restrictive binary. Are there not other entities which generate text? What if a gorilla uses ASL that is transcribed? ELIZA could generate text, after a fashion, as a precursor to LLM. It seems like there's a number of automated processes that could take data and generate text, sort of, like weather reports, no? So I think the only thing a mythical detector could determine would be LLM, or non-LLM, and let us take it from there. But detectors are bunk; I've had first-hand experience with that.
- hef19898 3y agoWe could combine those, couldn't we?
- yowlingcat 3y agoYou could but is there any reason to believe these two noisy signals wouldn't result in more combined noise than signal? Sure, it's theoretically possible to add two noisy signals that are uncorrelated and get noise reduction, but is it probable this would be such a case?
- cconstantine 3y agoYes, you can :) It all depends on the properties of the signal and the noise. In photography you can combine multiple noisy images to increase the signal to noise ratio. This works because the signal increases O(N) with the number of images but the noise only increases O(sqrt(N)). The result is that while both signal and noise are increasing, the signal is increasing faster. I have no idea if this idea could be used for AI detection, but it is possible to combine 2 noisy signals and get better SNR.
- NeoTar 3y agoIf the noisy signals are not completely correlated then the signal would be enhanced; however in this case I imagine that there is likely to be a strong correlation between different tools which would mean adding additional sources may not be so useful.
- TheSpiceIsLife 3y agoSome kind of Voigt-Kampff Test, perhaps.
- moffkalast 3y agoSomething something cells, interlinked.
- LordDragonfang 3y agoTBH, a properly-administered polygraph is probably more accurate than OpenAI's detector (of course, "properly administered" requires the subject to be cooperative and answer very simple yes or no questions, because a poly measures subconscious anxiety, not "truth")
- carapace 3y agoPolygraph is pseudo-science, it measures nothing.
- LordDragonfang 3y agoI mean, it literally and factually measures multiple your body's autonomous responses - all of which are provably correlated with stress. That's what a polygraph machine is. Saying it measures nothing is factually incorrect. You can't detect "truth" from that, but you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforward question in a situation where they don't have to be particularly nervous (which is why it's not very useful for interrogating a stressed criminal suspect, and should absolutely be inadmissible in court). But everyone knows that it's not very reliable in almost every circumstance it's used. My point is that while only marginally better than chance, it's still better than chance, unlike the OpenAI's detector, which is significant worse than chance.
- rndgermandude 3y agoRight. The point is: it absolutely does NOT measure what it claims to measure, i.e. truthfulness. You can detect indicators of stress... or hot weather... or stage-fright (admittedly a form of stress)... or too much caffeine... or an underlying (maybe undiagnosed) medical condition, etc. So it does not even necessarily measure "stress". It's about as useful as the so called "fruit machine" which they used to test for homosexuality[0], in that it is utterly useless while at the same time can be quite ruinous for people. People have been fired over polygraph "fails", and while not admissible in courts, people probably have been fingered for crimes after they failed polygraphs. Also, criminals have gone free after passing polygraphs[1]. >But everyone knows that it's not very reliable in almost every circumstance it's used. You and I may know that. But a lot of people actually do not. That's why it's still used. Either because people administering those tests think it's "good science", or because those people administering it know that while it's all bullshit the person they are testing might not know that and break down and admit to things. Remember that fake polygraph on the show The Wire, which was just a copier they strapped to the suspect. If I remember correctly that was based upon true events. A quick google shows e.g. you can hire "polygraphers" to e.g. "test" if your partner was unfaithful, making claims such as: "However, assuming that you have a good polygrapher with a fair amount of experience in working with betrayal trauma, you're going to get results that are at least 90% accurate or better."[2] The US (and probably a lot of other) government(s) like their polygraphs very much, too[3]. > you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforward question in a situation where they don't have to be particularly nervous Uhmm, if somebody sat me down in a room, strapped all kinds of "science" to my body and then asked me questions, I'd be quite nervous regardless of whether I am truthful or not. In fact, I'd be even more nervous knowing it's a polygraph and bullshit, because I cannot know if the person administrating it would know that too. If that somebody then asked me "Have you ever killed a prostitute?", or "Have you ever colluded with the enemy?", or "Have you ever cheated on your partner?", or "Have you ever stolen from your employer?", for example, my stress would certainly peak despite being able to confidently and truthfully answer "No!" to all of those questions. And I am sure the polygraph would "measure" my "stress". [0] Yes, that was a real thing too. https://en.wikipedia.org/wiki/Fruit_machine_(homosexuality_test) https://en.wikipedia.org/wiki/Fruit_machine_(homosexuality_t... [1] E.g. the Green River Killer Gary Ridgway passed a polygraph, so the police turned their resources to another suspect who failed the polygraph. That was in 1984. Ridgway remained free until his arrest in 2001. He killed at least 4 more times after the investigation stopped focusing on him after that "passed" polygraph. [2] https://www.affairrecovery.com/newsletter/founder/use-abuse-polygraph https://www.affairrecovery.com/newsletter/founder/use-abuse-... [3] https://support.clearancejobs.com/t/the-differences-between-counterintelligence-lifestyle-and-full-scope-polygraphs/46 https://support.clearancejobs.com/t/the-differences-between-...
- xattt 3y agoI see no reason why watermarking can’t be broken by having someone simply rephrase/redraw the output. Yes, it’s still work, but it’s one step removed from having to think up of the original content.
- whimsicalism 3y agoWatermarking was never going to be successful except for the most naive uses.
- SkyPuncher 3y agoIt can likely work in images where you can make subtle, human-undetectable tweaks across thousands/millions of pixels, each with many possible values. Nearly impossible across data with a couple hundred characters and dozens to thousands of tokens.
- whimsicalism 3y agoright but the non-naive approach would be to add noise or have a dumber model rewrite the image. agreed it is easier with images though
- BestGuess 3y agoTaking away tools don't seem to me like the best response same way taking away things tends never to be. If the problem is people not using it right, that seems to me like it would be designed wrong for what people need it for. Like if the issue is using it wrong with too little sentences, then put a minimum sentence or something to have that minimum likelihood. Same goes for representing what it means. If people don't understand statistics or math and such, then show what it means with circles or coins or stuff like that. Point is don't seem ever a good thing for options to get removed, especially if it's for bein cynical and judgin people like they're beneath deservin it. Don't make no sense.
- insanitybit 3y agoThe problem isn't people not using it right, the problem is that the tool can never work and just by being out in the world it would cause harm. If I have a tool that returns a random number between 0 and 1, indicating confidence that text is AI generated, is that tool good? Is it ethical to release it? I'd say no, it isn't. Removing the option is far better because the tool itself is harmful.
- CatWChainsaw 3y ago>just by being out in the world it would cause harm sounds like AI rather than AI detection to me. :)
- BestGuess 3y agoI don't agree with that premise. I don't know that it can't work, that'd suggest something like no matter what it's worse than a coin flip. I don't think it's that bad or at least nobody showed me anything of it being that bad. You'd have to show me that it can't work and that seems to me a pretty big ask I know
- insanitybit 3y agoAll that has to be shown is that the tool is as bad as or worse than random today, in order to remove it today.
- constantcrying 3y agoEven the idea of it is bad, ChatGPT is supposed to write indistinguishably from a human. The "detector" has extremely little information and the only somewhat reasonable criteria are things like style, where ChatGPT certainly has a particular, but by no means unique writing style. And as it gets better it will (by definition) be better at writing in more varied styles.
- Teever 3y agoNitpick: ChatGPR is supposed to write in a way that is indistinguishable from a human, to another human. That doesn't mean that it can't be distguishable by some other means.
- CookieCrisp 3y agoI think for small amounts of text there's no way around it being indistinguishable to a machine and not distinguishable to a human. There just aren't that many combinations of words that still flow well. Furthermore as more and more people use it I think we'll find some humans changing their speech patterns subconsciously more to mimic whatever it does. I imagine with longer text there will be things they'll be able to find, but, I think it will end up being trivial for others to detect what those changes are and then modifying the result enough to be undetectable.
- jerf 3y agoI think for this sort of problem it is more productive to think in terms of the amount of text necessary for detection, and how reliable such a detection would be, than a binary can/can't. I think similarly for how "photorealistic" a particular graphics tech is; many techs have already long passed the point where I can tell at 320x200 but they're not necessarily all there yet at 4K. LLMs clearly pass the single sentence test. If you generate far more text than their window, I'm pretty sure they'd clearly fail as they start getting repetitive or losing track of what they've written. In between, it varies depending on how much text you get to look at. A single paragraph is pretty darned hard. A full essay starts becoming something I'm more confident in my assessment. It's also worth reminding people that LLMs are more than just "ChatGPT in its standard form". As a human trying to do bot detection sometimes, I've noticed some tells in ChatGPT's "standard voice" which almost everyone is still using, but once people graduate from "Write a blog post about $TOPIC related to $LANGUAGE" to "Write a blog post about $TOPIC related to $LANGUAGE in the style of Ernest Hemmingway" in their prompts it's going to become very difficult to tell by style alone.
- specproc 3y agoI'm still interested in this line of enquiry. These models are clearly not good enough for decision-making, but still might tell an interesting story. Here's an easily testable exercise: get a load of news from somewhere like newsapi.ai, run it through an open model and there should be a clear discontinuity around ChatGPT launch. We can assume false positives and false negatives, but with a fat wadge of data we should still be able to discern trends. Certainly couldn't accuse a student of cheating with it, but maybe spot content farms.
- amelius 3y agoThey could certainly keep a database of things generated by /their/ AI ...
- andy99 3y ago> The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. I saw that this report came out today which frankly is baffling: https://gpai.ai/projects/responsible-ai/social-media-governance/Social%20Media%20Governance%20Project%20-%20July%202023.pdf https://gpai.ai/projects/responsible-ai/social-media-governa... (Foundation AI Models Need Detection Mechanisms as a Condition of Release [pdf])
- feoren 3y agoIndeed it's not possible. Say you had a classifier that detected whether a given text was AI generated or not. You can easily plug this classifier into the end of a generative network trying to fool it, and even backpropagate all the way from the yes/no output to the input layer of the generative network. Now you can easily generate text that fools that classifier. So such a model is doomed from the start, unless its parameters are a closely-guarded secret (and never leaked). Then it means it's foolable by those with access and nobody else. Which means there's a huge incentive for adversaries to make their own, etc. etc. until it's just a big arms race. It's clear the actual answer needs to be: we need better automated tools to detect quality content, whatever that might mean, whether written by a human or an AI. That would be a godsend. And if it turned into an arms race, the arms we're racing each other to build are just higher-quality content.
- tessierashpool 3y ago> You can easily plug this classifier into the end of a generative network trying to fool it, and even backpropagate all the way from the yes/no output to the input layer of the generative network. Now you can easily generate text that fools that classifier could you contextualize your use of the word "easily" here? I feel like "easily" might mean "with infinite funds and frictionless spherical developers."
- alexeldeib 3y agoGANs are established engineering. Infinite funds and frictionless spheres aside, you don't need to break ground, but copy/paste/glue existing code with some comprehension. LLMs are newer than GANs afaik, it just so happens GANs are a good fit here, not that one is "smarter" or "dumber".
- asddubs 3y agoThe whole problem with AI is that it's able to copy some of the superficial indicators of quality content while feeding you lies. You cannot detect quality content without detecting truthfulness. Any heuristic you use in place of that can be copied without actually providing value (which is exactly what ChatGPT does now, when it gets things wrong)
- itsmoneyzzzzu 3y ago[dead]
- airjtaiwjt 3y agoBased on my experience from grad school, I would bet plenty of the professors who fail students because ChatGPT said ChatGPT might have written something honestly don't care whether it's true or not, as long as it shifts liability away from themselves onto someone else
- burtonator 3y agoIt's really disturbing to me how many people don't realize it's not possible. It's like asking a 747 to be made into a dog. It's completely nonsensical to me.
- nradov 3y agoDr. Michio Kaku claimed in an interview that it may eventually be possible for quantum computers to guarantee a certain level of truthfulness. I didn't really follow his argument and it seemed a little hand-wavey, but I can't prove that he's wrong. https://mkaku.org/home/tag/quantum-computing/ https://mkaku.org/home/tag/quantum-computing/
- c-cube 3y agoNo need to prove! Any well calibrated bullshit detector should ring at 110dB when applied to claims mixing AI and quantum computing. (that said, "may eventually be possible" is so weak a claim it's already meaningless. Quantum fluctuations may eventually turn me into a potato but it's not keeping me up at night)
- nancyhn 3y agoI agree, transparency is essential, especially when it comes to AI applications. Many underestimate the complexity of distinguishing AI-written content from human-written, especially for short texts. There's a danger in trusting tools claiming to provide absolute certainty in this regard; no current technology can guarantee 100% accuracy. This incident underscores the need for a more realistic understanding of AI capabilities and limitations in text generation detection.
- Generalia 3y ago[dead]