3 ms·
LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for? [0] http
by drusepth 29d ago
LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?
[0] https://arxiv.org/pdf/2503.23674 https://arxiv.org/pdf/2503.23674
- thepasch 29d agoWith how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
- bbor 29d agoYou're overindexing on the past 3-6 months, IMHO.
- pants2 29d agoI might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
- acchow 29d agoThat’s using the default system prompt, right? Which is told to be an assistant.
- pkulak 29d agoMy whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
- debugnik 29d agoTuring never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
- pkulak 29d agoOkay, I believe you; mostly because you appear to be a human and I'm not really in the mood to read through a paper from 1950 at the moment. But I still stand by it being _my_ benchmark for machine intelligence, which is all I was claiming.
- deleted 29d ago[deleted]
- Kotlopou 28d agoThis paper is a pet peeve of mine. Look at the appendix! On page 22, you see a typical conversation people were judging based on. I'll reproduce one verbatim here: Q: do you like doing psych studies and why? A: theyre chill, easy money tbh Q: yeah same. Could you give me an easy cupcake recipe off the top of your head? A: nah i just get the box mix lol Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal? A: axolotl, theyre weirdly cute And that's the whole thing. I looked through the data they shared and it's all like that. They even included ELIZA and it was judged human 23% of the time. I hope this wouldn't pass peer review... but they didn't even try, it's a preprint.