3 ms·
If we could show the current models to someone like Alan Turing, I am sure he would conclude that we have AGI.
by kkoncevicius 2mo ago
If we could show the current models to someone like Alan Turing, I am sure he would conclude that we have AGI.
- stymaar 2mo agoAnd after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb
- klibertp 2mo agoI think of the Turing test as one of the starting lines, along with image recognition ("a summer break project for a group of grad students" resisted being solved for decades). It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
- stavros 2mo agoHow are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).
- ilaksh 2mo agoYes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.
- stavros 2mo agoThat tracks, I wonder what the people who say that the models are too dumb expect to see. Miracles?
- stymaar 2mo agoNot making mistakes my four years old would not. They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.
- nearbuy 2mo agoIt also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home https://www.trackingai.org/home) The test wasn't made to accurately measure IQs that high.
- hluska 2mo agoDumb is too simplistic, but they sure don’t pass for being human. That would be the intent of a Turing test… not IQ.
- ozgung 2mo agoI wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.
- hnbad 2mo agoYou need to be smarter (or rather: more knowledgeable in the problem domain) than the model to be able to use it efficiently. Hallucinations are still a problem occasionally but a bigger one is failure of imagination. Even Claude Fable lacks a holistic understanding of many domains it wasn't obviously trained on. The biggest problem with AI (if we assert that LLMs can be the basis of AI) is that these models will make mistakes that exist in an entirely different category of the kind of mistakes humans will make. As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
- kkoncevicius 2mo ago> much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied. This sounds interesting on it's own. I would be curious to hear more if you are willing to share.
- hnbad 1mo agoWhich part? I assume you mean the denial? A simple example would be "work to rule": in many professions work processes are heavily regulated (whether by law or by corporate guidelines) but the unspoken assumption is that you know which rules you should ignore and which ones you actually need to follow - but if you tried to find this out by asking "is this a rule I need to follow or not" you would get the clearly incorrect answer that all rules must be followed; of course if you did follow all the rules (aka "work to rule") you would be disciplined for failing to meet quotas (because you can't be disciplined for following the rules).
- txzl 2mo agoI just asked Fable 5 max to create a 2d game about caterpillar climbing a tree and eating fruits. The game looks good - animations, 8-bit aesthetics, procedural tree branching, but the tree's branches are dead ends. You can't go back once you started climbing a branch. Yes, LLM doesn't have a reliable way to test it's game yet. All the screenshots, and playwright tests will never be enough to test even a simple game. But can we call a machine doing such mistakes a general intelligence? It has no embodied intelligence. No way to experience time the way we do. All it has is text. Yes, they can have images, sound too, but no big models (at least those we are supposed to use for coding) currently are native with video as far as I am concerned. And I am not sure that just video without embodied experience is enough to understand the world the way humans do. Of course we can get incredible results from machines that have a very different experience of the world than we do. But is this a general intelligence? I guess "general" is supposed to mean being able to do everything any human can do (minus the skills requiring a body)?
- ipsod 2mo agoCurrent Gemini Flash models can take video input. They're not hyper-specialized coders, but they're better than the competition on many tasks. They seem to be better with spacial reasoning, as well - they are the best choice for OpenSCAD, for example.
- stymaar 2mo agoI think it's Karpathy who coined the term “jagged intelligence”. LLMs are both extraordinary smart in domain they have been explicitly trained on (like Math) and positively dumb on things they haven't.
- dotancohen 2mo agoI'm that way too. So are most people. That's why you shouldn't listen to your pop celebrities for political advice.
- Eggpants 2mo agoWould you also consider a database of questions and answers smart? LLM are basically lossy text compression databases with a clever query method. Useful for sure but it’s not thinking, it’s recall. Just look at some training sets to see how the sausage is made: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-80k-v1 https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-...
- jaccola 2mo agoI don’t think he would. The Turing Test as originally formulated is a bit ambiguous but by most non-incentivised interpretations LLMs do not pass it. In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.
- throwatdem12311 2mo agoYou say that they pass the Turing test yet every post on HN complains about the way LLMs write so clearly they haven’t passed it yet because we can still tell it’s a bot.
- hk__2 2mo agoI think we’re just adapting. LLM felt kind of magical at first, and now we’re all experts in detecting AI slope.
- pinkgolem 2mo agoYou can also change the writing style with a prompt quit easily, if the same person would publish millions of articles, we would also recognize them.
- PunchTornado 2mo agoI don't understand. do you think that people can't figure it out if they speak with llms or not from a few turns?
- quantumspandex 2mo agoWe don't train current LLMs to mimic the average human's writing style. We train it to be smart, helpful, and knowledgable, too knowledgeable for a human. We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.
- throwatdem12311 2mo ago> but it would just sound dumb or biased We have one of those: Grok.
- hobofan 2mo agoYes, and it's indistinguishable from the average X user.
- chippiewill 2mo agoTuring was a helluva smart guy, but he doesn't have the benefit of hindsight.
- deleted 2mo ago[deleted]