3 ms·
It hasn’t been passed and no one cares about it because it’s basically an end goal. No lab can hit it so they can’t juice the crazy Turing benchmark 3000 for ma
by jaccola 1mo ago
It hasn’t been passed and no one cares about it because it’s basically an end goal. No lab can hit it so they can’t juice the crazy Turing benchmark 3000 for marketing.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
- layla5alive 1mo agoThe test was not "after thousands of hours of conversing with them, knowing they're AI, THEN see if you can tell them apart blindly." Were 2010 you to be in a real turing test with an arbitrary erudite human and a 2026 frontier LLM, not knowing LLMs existed, you'd probably struggle
- smohare 1mo agoI doubt this entirely. It might be quite difficult for said human to discern whether a simple passage were generated sans such accumulated experience in reading AI text, true. But LLMs do not converse like humans in ways that have always been essentially immediately obvious.
- zoho_seni 1mo agoHave you seen how many people talk to bots these days thinking is a real person. Or that are even in a relationship with them or friends.
- jurgenburgen 1mo agoSometimes I wonder if LLMs are just revealing a section of the population with untreated mental illness or if LLMs are actively exacerbating mental illness. We might eventually regret exposing the general population to such a new technology without almost any safeguards.
- customguy 1mo agoSo? The people who can do it prove it can be done - the people who can't don't prove the opposite. Might as well claim all math is wrong because most people don't understand it.
- eloisius 1mo agoHave you seen how many people form one-sided bonds with stars that don’t know they exist? Lonely people suspend disbelief to find some comfort. It’s not proof that the chatbot is indistinguishable from a human companion.
- zahlman 1mo agoOkay, but the same could have been said about ELIZA.
- jaccola 1mo agoThis is always the most silly argument. The original test was ambiguous but for sure the human was trying to prove themselves human. So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this! Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
- akoboldfrying 1mo ago> It hasn’t been passed https://arxiv.org/abs/2503.23674 https://arxiv.org/abs/2503.23674 From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
- jaccola 1mo agoLow n, time bound, not reproduced. And look at their example conversations… And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
- akoboldfrying 1mo ago> Low n Are you serious? From the paper: > We recruited 126 participants from the UCSD psychology undergraduate subject pool and 158 participants from Prolific (Prolific, 2025). Each human participated in 8 rounds. > time bound The time bound of 5 minutes was suggested by Turing himself in his original paper. > not reproduced It was reproduced across two populations within the paper. > And look at their example conversations This is irrelevant.
- jaccola 1mo agoA paper can’t reproduce its self. And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly. And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition). And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time. The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!