3 ms·
Low n, time bound, not reproduced. And look at their example conversations… And people forget that sometimes humans message twice. An LLM can only respond. So
by jaccola 1mo ago
Low n, time bound, not reproduced. And look at their example conversations…
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
- akoboldfrying 1mo ago> Low n Are you serious? From the paper: > We recruited 126 participants from the UCSD psychology undergraduate subject pool and 158 participants from Prolific (Prolific, 2025). Each human participated in 8 rounds. > time bound The time bound of 5 minutes was suggested by Turing himself in his original paper. > not reproduced It was reproduced across two populations within the paper. > And look at their example conversations This is irrelevant.
- jaccola 1mo agoA paper can’t reproduce its self. And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly. And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition). And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time. The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!
- akoboldfrying 1mo ago> And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly. Right, the test duration was left unspecified. This means any duration is acceptable. Including, for example, the only duration actually mentioned by Turing himself in his paper. Or do you have a more authoritative source on which durations are acceptable? > If one human on earth can consistently get it right then it hasn’t been passed Says who? Not Turing. Probably he didn't say that because it would make the test both impractical and overly conservative. > The fact these researchers have to keep adding bounds What "bounds"? The speed of the goalposts here is just amazing.