4 ms·
If you were trying to predict the direction a stock will move (up or down) and it was right 99.9% of the time, would you use it or not?
by dontreact 3y ago
If you were trying to predict the direction a stock will move (up or down) and it was right 99.9% of the time, would you use it or not?
- a13o 3y agoThis is a strawman. First, the AI detection algorithms can't offer anything close to 99.9%. Second, your scenario doesn't analyze another human and issue judgement, as the AI detection algorithms do. When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.
- dontreact 3y agoIt's not a strawman. There are many fundamentally unpredictable things where we can't make the benchmark be 100% accuracy. To make it more concrete on work I am very familiar with: breast cancer screening. If you had a model that outperformed human radiologists at predicting whether there is pathology confirmed cancer within 1 year, but the accuracy was not 100%, would you want to use that model or not?
- frumper 3y agoIt's a strawman because they aren't comparable to AI detection tests. A screening coming back as possible cancer will lead to follow up tests to confirm, or rule out. An AI detection test coming back as positive can't be refuted or further tested with any level of accuracy. It's a completely unverifiable test with a low accuracy.
- dontreact 3y agoYou are moving the goalposts here. The original claim I am responding to is "A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process." I agree that there are places where we shouldn't put AI and that checking whether something is an LLM or not is one of them. However I think the sentence above takes it way too far and breast cancer screening is a pretty clear example of somewhere we should accept AI even if it can sometimes make mistakes.
- frumper 3y agoThe thread is about tools to evaluate LLMs. Please re-read my comment in that light and generously assume I'm talking about that.
- skissane 3y ago> Second, your scenario doesn't analyze another human and issue judgement, as the AI detection algorithms do. > When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc. Is the problem here the algorithms or how people choose to use them? There’s a big difference between treating the results of an AI algorithm as infallible, and treating it as just one piece of probabilistic evidence, to be combined with others, to produce a probabilistic conclusion. “AI detector says AI wrote student’s essay, therefore it must be true, so let’s fail/expel/etc them” vs “AI detector says AI wrote student’s essay, plus I have other independent reasons to suspect that, so I’m going to investigate the matter further”
- a13o 3y agoThat's exactly why the stock analogy doesn't work. People don't buy algorithms, they buy products - such as detectors or predictors. You necessarily have to sell judgement alongside the algorithm. So debating the merits of an algorithm in a vacuum, when the issue being raised is the human harm caused by detector products, is the strawman.
- skissane 3y ago> People don't buy algorithms, they buy products - such as detectors or predictors. You necessarily have to sell judgement alongside the algorithm. Two people can buy the same product yet use it in very different ways: some educators take the output of anti-cheating software with a grain of salt, others treat it as infallible gospel. Neither approach is determined by the product design in itself, rather by the broader business context (sales, marketing, education, training, implementation), and even factors entirely external to the vendor (differences in professional culture among educational institutions/systems).