5 ms·
I for one value the transparency of being able to tell human-written content apart, especially when moving into a time where we’re suddenly questioning an AI ag
by 9dev 2mo ago
I for one value the transparency of being able to tell human-written content apart, especially when moving into a time where we’re suddenly questioning an AI agent's motivations. If that is ingrained so deeply into these systems they cannot rip it out without us noticing, there’s at least one more safeguard. I know how ludicrous this sounds, but you can’t deny AI development is speeding up to scary levels.
Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
- josephg 2mo agoMe too! And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students. This article says any AI watermarking can be defeated trivially, by running the text through a tool which strips weird Unicode characters (fine) and then runs it through an LLM which replaces all words with similar words (whaaattt?). Running text through another - probably much lower quality llm which scrambles the words you use would make slop even sloppier. And it’s another whole step you have to know to do. A great many people who use LLMs to avoid doing work will not know about these extra steps, or not make use of these sort of tools.
- sib 2mo ago>> And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students. I would say that's not necessarily the case unless there are zero false positives. In fact, your university situation is exactly where a detection system that works 50-90% of the time would be a nightmare if a meaningful share of the 10-50% errors were false positives.
- josephg 2mo agoA watermarking system like this should have an incredibly low false positive rate. Orders of magnitude lower than the current crop of AI detector tools.
- nonethewiser 2mo agoAka some kids are gonna get majorly fucked. Just not many.
- josephg 2mo agoThe lazy kids, yeah. But, it was always the lazy kids who are cheating the most with LLMs. EDIT: Sorry, I misread your comment above. Yeah, hopefully orders of magnitude fewer kids than the number who are getting falsely accused of cheating with LLMs now. Increasing the accuracy of these systems - both in terms of false positives and false negatives - seems like a good thing.
- nonethewiser 2mo agoLazy? If some kids are false positives (detected as using ai but didn’t)then how are they lazy?
- rightbyte 2mo agoApple's CSAM hash filter removing peoples photos (did it notify the police too?) could be an indicator of how false positives on big scale work out.
- josephg 2mo agoHow accurate is apple's CSAM scanner? 90%? 99%? A proper watermarking system should be able to have an arbitrarily high accuracy - as many nines as you want. And it should be able to actually report the accuracy of its judgements. If you're worried about kids being falsely accused of cheating using LLMs, you should be cheering on these developments.
- nonethewiser 2mo agoAnd students could just verify their ai generated code doesnt have the mark. Im very confused how this is even supposed to work at face value. 1) If the verification can be done by anyone, then anyone can bypass it. 2) If it can only be done by anthropic then the government or whomever has the special privilege (not everyone otherwise this is just #1) has to make a specific request Is the point of the legislation to accurately classify text in general or to simply detect the true positives?
- nradov 2mo agoEven if some of the commercial LLMs embed text watermarks, students will still be able to use open source LLMs with no watermarks. Universities need to stop goofing around and restructure every part of their curriculum around the assumption that students will use those tools. Trying to treat it as cheating is just pissing into the wind. Circumstances have fundamentally changed and there's no going back.
- josephg 2mo agoAh yes, cheating is really the fault of universities for not fundamentally restructuring their courses more. Why not both? Stenographically fingerprint llm output to make cheating risky. And use other approaches too. Most people have no idea open source LLMs exist, let alone how to use them. Right now, they’re much worse than the frontier models.
- vaylian 2mo ago> Me too! And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students. What about false positives? Imagine being a honest student and then the software declares your work to be AI-generated. How do you defend against that claim? The detection software is a black box and is likely running as a cloud service, so that you have no realistic options for reverse-engineering the false positive detection.
- jghn 2mo agoThe problem is that the false positive rate needs to be exactly 0.00% and that's not going to happen.
- nonethewiser 2mo agoThis doesn’t add up. How do you think you will know the text was AI generated but the person trying to deceive you won’t be able to undo the watermark?
- optionalsquid 2mo agoPractically speaking, most people probably won't know how to undo the watermark. Especially when the suggested method is to run the text through a local model
- nonethewiser 2mo agoDoesnt matter - it ruins integrity and the people actively trying to deceive you wont be prevented.
- optionalsquid 2mo agoWhat do you mean that it ruins integrity?
- nonethewiser 2mo agoI mean you cant be sure the text isnt AI.
- cwillu 2mo agoMost people are entirely capable of copy/pasting their text into the inevitable dozens of web-based tools of various levels of sketch to do it for them.
- potsandpans 2mo agoNow you just get to think, "did this person strip the stenography from the content?" Nothing will save us. You can't automate trust.
- nradov 2mo agoI deny that there is anything scary about AI development. It's working fine for me so far.