3 ms·
Please report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me imme
by Noumenon72 1mo ago
Please report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.
- dozerly 1mo agoYea, I did two and then harrumphed in annoyance that I was expected to do all 10.
- DonHopkins 1mo agoWe must do something about this immediately! Immediately! Immediately! https://www.youtube.com/watch?v=jLO7VrRij_M https://www.youtube.com/watch?v=jLO7VrRij_M
- stranded22 1mo agoYes. Did one - saw that I wouldn’t get feedback until I have completed all 10 (if at all) and noped out.
- StilesCrisis 1mo agoI did five, then gave up and just pressed A until I reached the end. I got 3/5 right.
- neoncontrails 1mo agoExactly the same here. 4/5.
- andai 1mo ago> Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses. Wouldn't this make it a worse measurement?
- neoncontrails 1mo agoWell, the alternative is a lot of us spammed A to get to the end, so the data quality is already horrendous.
- Noumenon72 1mo agoI don't want to know whether Claude can pass the Turing test of someone who has never seen Claude's style before. I want to know whether I will start detecting the watermark everywhere and come to loathe its taint, and I want to learn how to identify it so I don't get tricked. Not interested in doing anything extra to help the website.
- qarl2 1mo agoI believe this is to test the theory that people can detect watermarked text. Teaching you how to identify watermarked text while the experiment is running would ruin the data.
- Phemist 1mo agoMaybe detecting watermarked text is a skill to be attained. Not allowing proper feedback and training will not allow people to notice the difference on time, thus ruining the data? Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.
- ipaddr 1mo agoDetecting watermarked text is hard. Detecting AI garbage text is easy but then classfiying the garbage further is beyond us.
- thaumasiotes 1mo ago> I believe this is to test the theory that people can detect watermarked text. > Teaching you how to identify watermarked text while the experiment is running would ruin the data. But that's complete nonsense. If it's possible to teach someone how to identify watermarked text, then you've already proven that people can detect watermarked text.
- barnabee 1mo agoI managed 0/10 - far worse than random chance Not sure what that says about me or the LLM but I guess I shouldn't worry too much about watermarking ruining the outputs…