6 ms·
We Have Made No Progress Toward AGI
- thisisnotauser 1y agoImma be honest with you, this is exactly his I would do that math, and that is exactly the lie I would tell if you asked me to explain it. This is me-level agi.
- x187463 1y ago> Which means these LLM architectures will not be producing groundbreaking novel theories in science and technology. Is it not possible that new theories and breakthroughs could result from this so-called statistical pattern matching? The information necessary could be present in the training data and the relationship simply never before considered by a human. We may not be on a path to AGI, but it seems premature to claim LLMs are fundamentally incapable of such contributions to knowledge. In fact, it seems that these AI labs are leaning in such a direction. Keep producing better LLMs until the LLM can make contributions that drive the field forward.
- 13years 1y agoCertainly random chance exists for discovery. But most revolutionary type discoveries come from deep understanding of the context. The contribution of LLMs to knowledge is more like that of search engines. It is still the human which possesses understanding that ultimately will be the principle source of innovation. The LLM can assist with navigating and exploring existing information. However, LLMs have significant downsides in this regard too. The hallucination problem is no joke. It can often mislead you and cause a loss of time on some tasks. Overall, they will be somewhat useful in some manner, but substantially less so than the present hype machine suggests.
- thefounder 1y agoThey seem precisely that: search engines. Instead to give you a list of webpages with possible answers they actually synthesise the results. A more direct analogy is the case where ChatGPT provides you two possible answers. Of course it could provide you more just like search engines provide more links.
- ebiester 1y agoThe hard part is that for all the things that the author says disprove LLMs are intelligent are failings for humans too. * Humans tell you how they think, but it seemingly is not how they really think. * Humans tell you repeatedly they used a tool, but they did it another way. * Humans tell you facts they believe to be true but are false. * Humans often need to be verified by another human and should not be trusted. * Humans are extraordinarily hard to align. While I am sympathetic to the argument, and I agree that machines aligned on their own goals over a longer timeframe is still science fiction, I think this particular argument fails. GPT o3 is a better writer than most high school students at the time of graduation. GPT o3 is a better researcher than most high school students at the time of graduation. GPT o3 is a better lots of things than any high school student at the time of graduation. It is a better coder than the vast majority of first semester computer science students. The original Turing test has been shattered. We're building progressively harder standards to get to what is human intelligence and as we find another one, we are quickly achieving it. The gap is elsewhere: look at Devin as to the limitation. Its ability to follow its own goal plans is the next frontier and maybe we don't want to solve that problem yet. What if we just decide not to solve that particular problem and lean further into the cyborg model? We don't need them to replace humans - we need them to integrate with humans.
- 13years 1y ago> GPT o3 is a better writer than most high school students at the time of graduation. All of these claims, based on benchmarks, don't hold up in the real world on real world tasks. Which is strongly supportive of the statistical model. It will be capable of answering patterns extensively trained on. But is quickly breaks down when you step outside that distribution. o3 is also a significant hallucinator. I spent quite a bit of time with it last weekend and found it to be probably far worse than any of the other top models. The catch is that it its hallucinations are quite sophisticated. Unless you are using it on material for which you are extremely knowledgeable, you won't know. LLMs are probability machines. Which means they will mostly produce content that aligns to the common distribution of data. They don't analyze what is correct, but only what is probable completions for your text by common word distributions. But when scaled to incomprehensible scales of combinatorial patterns, it does create a convincing mimic of intelligence and it does have its uses. But importantly, it diverges from the behaviors we would see in true intelligence in ways that make it inadequate for solving many of the kinds of tasks we are hoping to apply them to. The being namely the significant unpredictable behaviors. There is just no way to know what type of query/prompt will result in operating over concepts outside the training set.
- thefounder 1y agoSo the “reasoning” text of openAI is no more than old broken Windows “loading” animation.
- xg15 1y agoI think there is still a widespread confusion between two slightly different concepts that the author also tripped over. If you ask an LLM a question, then get the answer and then ask how it got to that answer, it will make stuff up - because it literally can't do otherwise: There is no hidden memory space in which the LLM could do its calculations, and also record which calculations it did, that it could then consult to answer the second question. All there is are the tokens. However if you tell the model to "think step by step", I.e. first make a number of small inferences, then use those to derive the final answer, you should (at least in theory) get a high-level description of the actual reasoning process, because the model will use the tokens of its intermediate results to generate the features for the final result. So "explain how you did it" will give you bullshit, but "think step by step" should work. And as far as my understanding goes, the "reasoning models" are essentially just heavily fine tuned for step-by-step reasoning.
- 13years 1y ago> that the author also tripped over The evidence for unfaithful reasoning comes from Anthropic. It is in their system card and this Anthropic paper. https://assets.anthropic.com/m/71876fabef0f0ed4/original/reasoning_models_paper.pdf https://assets.anthropic.com/m/71876fabef0f0ed4/original/rea...
- advisedwang 1y agoMy understanding was that chain-of-thought is used precisely BECAUSE it doesn't reproduce the same logic that simply asking the question directly does. In "fabricating" an explanation for what it might have done if asked the question directly, it has actually produced correct reasoning. Therefore you can ask the chain-of-thought question to get a better result than asking the question directly. I'd love to see the multiplication accuracy chart from https://www.mindprison.cc/p/why-llms-dont-ask-for-calculators https://www.mindprison.cc/p/why-llms-dont-ask-for-calculator... with the output from a chain-of-thought prompt.
- deleted 1y ago[deleted]
- mark_l_watson 1y agoI mildly disagree with the author, but would be happy arguing his side also on some of his points: Last September I used ChatGPT, Gemini, and Claude in combination to write a complex piece of code from scratch. It took four hours and I had to be very actively involved. A week ago o3 solved it on its own, at least the Python version ran as-is, but the Common Lisp version needed some tweaking (maybe 5 minutes of my time). This is exponential improvement and it is not so much the base LLMs getting better, rather it is: familiarity with me (chat history) and much better tool use. I may be be incorrect, but I think improvements in very long user event and interaction context, increasingly intelligent tool use, perhaps some form of RL to develop per-user policies for improving incorrect tool use, and increasingly good base LLMs will get us to a place that in the domain of digital knowledge work where we will have personal agents that are AGI for a huge range of use cases.
- 13years 1y ago> where we will have personal agents that are AGI for a huge range of use cases We are already there for internet social media bots. I think the issue here is being able to discern the correct use cases. What is your error tolerance? For social media bots, it really doesn't matter so much. However, mission critical business automation is another story. We need to better understand the nature of these tools. The most difficult problem is that there is no clear line for the point of failure. You don't know when you have drifted outside of the training set competency. The tool can't tell you what it is not good at. It can't tell you what it does not know. This limits its applicability for hands-off automation tasks. If you have a task that must always succeed, there must be human review for whatever is assigned to the LLM.
- xandrius 1y agoDepends on what AGI means to you. If writing a program that you find complicated counts then it is. I do think that writing code is a specific type of task that a statistical machine can do well without actually bringing us closer to AGI.
- setnone 1y ago> All of the current architectures are simply brute-force pattern matching This explains hallucinations and i agree with 'braindead' argument. To move toward AGI i believe there should be some kind of social awareness component added which is an important part of human intelligence.
- moralestapia 1y agoI really dislike what I now call the American We. "We made it!" "We failed!" written by somebody who doesn't have the slightest connection to the projects they're talking about. e.g. this piece doesn't even have an author but I highly doubt he has done anything more than using chatgpt.com a couple times. Maybe this could be the Neumann's law of headlines: if it starts with We, it's bullshit.
- alehlopeh 1y agoI’ve been saying this for ages. People use “we” way too freely.
- lgas 1y agoWe're sorry, we'll try to do better.
- echoangle 1y agoIsn’t the „we“ supposed to mean „humanity“?
- jibal 1y ago"I highly doubt he has done anything more than using chatgpt.com a couple times." Based on what, other than BAD (being a you-know-what)? And intelligent people understand that "we" to refer to humans generally.
- jug 1y agoRed flag nowadays is when a blog post tries to judge whether AI is AGI. Because these goal posts are constantly moving and there is no agreed upon benchmark to meet. More often than not, they reason why exactly something is not AGI yet from their perspective, while another user happily use AI as a full-fledged employee depending on use case. I’m personally using AI as a coding companion and it seems to be doing extremely well for being brain dead at least.
- DrFalkyn 1y agoIt’s AGI when it can improve itself with no more than a little human interaction
- echoangle 1y agoWho is using AI as full-fledged employees?
- goatlover 1y agoHas the AI replaced you yet? Or is it a productive tool you make use of it? It's the difference between the Enterprise Computer and Lieutenant Commander Data.
- Sohcahtoa82 1y ago> these goal posts are constantly moving No, they're not. The only people who have made claims about achieving AGI are grifters trying to stir up hype or funding. Their opinions are not important.
- jibal 1y ago[flagged]
- tboyd47 1y agoFascinating look at how AI actually reasons. I think it's pretty close to how the average human reasons. But he's right that the efficiency of AI is much worse, and that matters, too. Great read.
- hnpolicestate 1y agoOne point that I think seperates AI and human intelligence is LLM's inability to tell me how it feels or it's individual opinion on things. I think to be considered alive you have to have an opinion on things.
- maebert 1y agoauthor says we made no progress towards agi, also gives no definition for what the "i" in agi is, or how we would measure meaningful progress in this direction. in a somewhat ironic twist, it seems like the authors internal definition for "intelligence" fits much closer with 1950s. good old-fashioned AI, doing proper logic and algebra. literally all the progress we made in ai in the last 20 years in ai is precisely because we abandoned this narrow-minded definition of intelligence. Maybe I'm a grumpy old fart but none of these are new arguments. Philosophy of mind has an amazingly deep and colorful wealth of insights in this matter, and I don't know why this is not required reading for anyone writing a blog on ai.
- 13years 1y ago> or how we would measure meaningful progress in this direction. "First, we should measure is the ratio of capability against the quantity of data and training effort. Capability rising while data and training effort are falling would be the interesting signal that we are making progress without simply brute-forcing the result. The second signal for intelligence would be no modal collapse in a closed system. It is known that LLMs will suffer from model collapse in a closed system where they train on their own data."
- maebert 1y agoI agree that those both are very helpful metrics, but they are not a definition of intelligence. yes, humans can learn to comprehend and speak language with magnitudes less examples than llms, however we also have very specific hardware for that evolved over millions of years — it's plausible that language acquisition in humans is more akin to fine-tuning in llms than training them from ground up. Either way, this metric is comparing apples to oranges when it comes to comparing real and artificial intelligence. model collapse is a problem in ai that needs to be solved, and maybe it's even a necessary condition for true intelligence, though certainly not a sufficient one, and hence not an equivalent definition of intelligence either.
- 13years 1y ago
- xg15 1y agoPeople ditch symbolic reasoning for statistical models, then are surprised when the model does, in fact, use statistical features and not symbolic reasoning.
- 13years 1y agoI think it is actually worse than that. The hype labs are still defiantly trying to convince us that somehow merely scaling statistics will lead to the emergence of true intelligence. They haven't reached the point of being "surprised" as of yet.
- mlsu 1y agoThis thing where AI can improve itself seems to me in violation of the second law. I'm not a physicist by training merely an engineer but my argument is as follows: - I think the reason humans are clever is because nature spent 6 billion years * millions of energetic lifetimes (that is, something on the order of quettajoules of energy) optimizing us to be clever. - Life is a system, which does nothing more than optimize and pass on information. An organism is a thing which reproduces itself, well enough to pass its DNA (aka. information) along. In some sense, it is a gigantic heat engine which exploits the energy gradient to organize itself, in the manner of a dissipative structure [1] - Think of how "AI" was invented: all of these geometric intuitions we have about deep learning, all of the cleverness we use, to imagine how backpropagation works and invent new thinking machines. All of the cleverness humanity has used to create the training dataset for these machines. This cleverness. It could not arise spontaneously, instead, it arose as a byproduct, from the long existence of a terawatt energy gradient. This captured energy was expended, to compress information/energy from the physical world, in a process which created highly organized structures (human brains) that are capable of being clever. - The cleverness of human beings and the machines they make is, in fact, nothing more than the byproduct of an elaborate dissipative structure whose emergence and continued organization requires enormous amounts of physical energy: 1-2% of all solar radiation hitting earth (terawatts), times 3 billion years (existence of photosynthesis). - If you look at it this way it's incredibly clear that the remarkable cleverness of these machines is nothing more than a bounded image, of the cleverness of human beings. We have a long way to go, before we are training artificial neural networks, with energy on the order of 10^30 joule [2]. Until then, we will not become capable of making machines that are cleverer than human beings. - Perhaps we could make a machine that is cleverer than one single human. But we will never have an AI that is more clever than a collection of us, because the thing itself must be, in a 2nd law sense, less clever than us, for the simple reason that we have used our cleverness to create it. - That is to say that there is no free lunch. A "superhuman" AI will not happen in 10, 100, or even 1,000 years, unless we find the vast amount of energy (10^30J) which will be required to train it. Humans will always be better and smarter. We have had 3 billion years of photosynthesis, this thing was trained in what, 120 days? A petajoule? [1] https://pmc.ncbi.nlm.nih.gov/articles/PMC7712552/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7712552/ [2] Where do we get 10^30J? Total energy hitting earth in one year: 5.5×10^24 J Fraction of that energy used by all plants: 0.05% Time plants have been alive on earth: 3 billion years You get to 8*10^30 if you multiply these numbers. Round down.
- nsonha 1y agoSo? Who even wants it? Whatever the definition is, sounds like AGI and sentient AI are really close concepts. Sentient AI is like a can of worms for ethics. On the other hands, while definitely not having AGI, we have all these building blocks for AI tools for decades to come, to build on top. We've only barely scratched the surface of it.