3 ms·
> They give useful probabilities Yes, compared to an all-or-nothing approach, it's better to be upfront about the uncertainty, especially if the tool surfaces
by berdario 2y ago
> They give useful probabilities
Yes, compared to an all-or-nothing approach, it's better to be upfront about the uncertainty, especially if the tool surfaces the probability by-sentence.
But how are those probabilities computed? You mention gptzero, but https://gptzero.me/technology https://gptzero.me/technology doesn't clarify at all how it works. They link papers using GPTZero (i.e. from other researchers), e.g. https://arxiv.org/pdf/2310.13606 https://arxiv.org/pdf/2310.13606
And these very same papers highlight how everything is still unknown
> Despite their wide use and support
of non-English languages, the extent of their zero-
shot multilingual and cross-lingual proficiency in
detecting MGT remains unknown. The training
methodologies, weight parameters, and the spe-
cific data used for these detectors remain undis-
closed.
GPTZero seems to be better than some of the alternatives, but other discussions here on HN when it was launched highlight all of the false positives and false negatives it yielded:
https://news.ycombinator.com/item?id=34556681 https://news.ycombinator.com/item?id=34556681
https://news.ycombinator.com/item?id=34859348 https://news.ycombinator.com/item?id=34859348
But all of that is pretty old, there have been a couple of posts in the last year about it, but both are about the business, rather than the quality of the tool itself.
https://hn.algolia.com/?dateRange=pastYear&page=0&prefix=false&query=gptzero&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=pastYear&page=0&prefix=fal...
So, to check if now it's any better I tried it myself: I got it to yield a false negative (50% human and 50% AI rating, for a text which was wholly AI-generated), and I haven't got it to yield a false positive.
But all of this is just anecdotal evidence, I haven't run a rigorous study.
For sure, if some competent people believe that the tool won't generate false positives, I'll be mindful of it and (In the rare cases in which I write a long posts/blog articles, etc.) I'll check that it doesn't erroneously flag what I write.
It's bittersweet: if a tool that can be relied upon really exist, that would be good news. But if that tool is closed source (just like ChatGPT, Gemini, etc.) that doesn't inspire confidence. What if the closed source detection tool will suddenly start erroneously flagging a subset of human texts which it didn't before?
At least, even with the closed source LLMs, we have a bunch of papers that explain their mechanism. I hope that GPTZero will be more forthcoming about the way it works.