4 ms·
Criticisms like this are levied against an excessively narrow (obsolete?) characterisation of what is happening in the AI space currently. After reading about
by pnut 2y ago
Criticisms like this are levied against an excessively narrow (obsolete?) characterisation of what is happening in the AI space currently.
After reading about o3's performance on ARC-AGI, I strongly suspect people will not be so flippantly dismissive of the inherent limits of these technologies by this time next year. I'm genuinely surprised at how myopic HN commentary is on this topic in general. Maybe because the implications are almost unthinkably profound.
Anyway, OpenAI, Anthropic, Meta, and everyone else are well aware of these types of criticisms, and are making significant, measurable progress towards architecturally solving the deficiencies.
https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough
- jokethrowaway 2y agoNah, the trick with o3 solving IQ tests seems to be that they bruteforce solutions and then pick the best option. That's why calls that are trivial for humans end up costing a lot. It still can't think and it won't think. LANGUAGE models (keyword: language) is a language model, it should be paired with a reasoning engine to translate the inner thought of the machine into human language. It should not be the source of decisions because it sucks at doing so, even though the network can exhibit some intelligence. We will never have AGI with just a language model. That said, most jobs people do are still at risk, even with chatgpt-3.5 (especially outside of knowledge work, where difficult decisions need to be taken). So we'll see the problems with AGI and the job market way earlier than AGI, as soon as we apply robotics and vision models + chatgpt 3.5 level intelligence. Goodbye baristas, goodbye people working in factories. Let's start working on a reasoning engine so we can replace those pesky knowledge workers too.
- esafak 2y agoThe important thing is that you can use inference-time computation to improve results. Now the race is on to optimize that.
- ithkuil 2y agoHow many attempts can you have when running an evaluation run of an ARC competition?
- sunnybeetroot 2y agoWe’ve had coffee machines that can make a perfect coffee with a touch of a button for at least a decade. How does GPT3.5 remove baristas given they could have already been removed?
- rafaelmn 2y agoReading the o1 announcement you could have been saying the same thing a year ago yet it's worse than Claude in practice and if it was all that's available - I wouldn't even use it if it was free - it's that bad. If OpenAI has demonstrated one thing is that they are a hype production machine and they are probably getting ready for next round of investment. I wouldn't be surprised if this model was equally useless as o1 when you factor in performance and price. At this point they are completely untrustworthy and untill something lands publicly for me to test it's safe to ignore their PR as complete BS.
- benterix 2y ago> yet it's worse than Claude in practice For most tasks - but not all. I normally paste my prompt in both and while Claude is generally superior in most aspects, there are tasks at which o1 performed slightly better.
- portaouflop 2y agoAny day now!
- wavemode 2y agoAGI is the new nuclear fusion.
- voidfunc 2y agoExcept AI is actually delivering value and on the path to AGI.. and nuclear fusion continues to be a physicists and engineering pipe dream.
- JasserInicide 2y agoWhat actual widespread non-shareholder value has AI given us?
- danielbln 2y agoConsidering the advancements we've seen in the last three years, this dismissive comment feels misplaced.
- portaouflop 2y agoLet’s just wait and see - what good can come from endless speculation and what ifs and trying to predict the future?
- gtirloni 2y agoConsidering the advancements we've seen in the last one year, it does not.
- mbesto 2y ago> I strongly suspect people will not be so flippantly dismissive of the inherent limits of these technologies by this time next year. People are flippantly dismissive of the inherent limits because there ARE inherent limitations of the technology. > Maybe because the implications are almost unthinkably profound. Maybe because the stuff you're pointing to are just benchmarks and the definitions around things like AGI are flawed (and the goalposts are constantly moving, just like the definition of autonomous driving). I use LLMs roughly 20-30x a day - they're an absolutely wonderful tool and work like magic, but they are flawed for some very fundamental reasons.
- qup 2y ago[flagged]
- greentxt 2y agoHumans are not flawed? Are robotaxi's not autonomous driving? (Could an LLM have written this post?)
- manquer 2y agoHumans are not machines , they have both rights that machines do not have and also responsibilities and consequences that machines will not have, for example bad driving will cost you money, injury , prison time or even death. Therefore AI has to be much better than humans at the task to be considered ready to be a replacement. —— Today robot taxis can only work in fair weather conditions in locations that are planned cities. No autonomous driving system can drive in Nigeria or India or even many european cities that were never designed for cars any time soon . Working in very specific scenarios is useful , but hardly measure of their intelligence or candidate for replacing humans for the task
- space_fountain 2y agoI hear people say this kind of thing but it confuses me. 1. What does inherit limitations mean? 2. How do we know something is an inherit limitation 3. Is it a problem if arguments for a particular inherit limitation also apply to humans? From what I’ve seen people will often say things like AI can’t be creative because it’s just a statistical machine, but humans are also “just” statistical machines. People might mean something like humans are more grounded because humans react not just to how the world already works but how the world reacts to actions they take, but this difference misunderstands how LLMs are trained. Like humans LLMs get most of their training from observing the world, but LLMs are also trained with re-enforcement learning and this will surely be an active area of research.
- netdevphoenix 2y agoYou remember when Google was scared to release LLMs? You remember that Googler that got fired because he thought the LLM was sentient? There is likely a couple of surprised still left in LLMs but no one should think that any present technology in its current state or architecture will get us to AGI or anything that remotely resembles it.
- gosub100 2y ago> Maybe because the implications are almost unthinkably profound. laundering stolen IP from actual human artists and researchers, extinguishing jobs, deflecting responsibility for disasters. yeah, I can't wait for these "profound implications" to come to fruition!
- lgas 2y agoThe implications of the technology are not impacted by how the technology was created or where the IP was sourced.
- formerly_proven 2y agoIt doesn’t really matter. “It works and is cost/resource-effective at being an AGI” is a fundamentally uninteresting proposition because we’re done at that point. It’s like debating how we’re going to deal with the demise of our star; we won’t, because we can’t.
- fabianhjr 2y ago> The question of whether a computer can think is no more interesting than the question of whether a submarine can swim. ~ Edsger W. Dijkstra LLMs / Generative Models can have a profound societal and economic impact without being intelligent. The obsession with intelligence only make their use haphazard and dangerous. It is a good thing court of laws have established precedent that organizations deploying LLM chatbots are responsible for their output (Eg, Air Canada LLM chatbot promising a non-existent discount being responsibility of Air Canada) Also most automation has been happening without LLMs/Generative Models. Things like better vision systems have had an enormous impact with industrial automation and QA.
- agentultra 2y agoThe conclusion of the article admits that in areas where stochastic outputs are expected these AI models will continue to be useful. It’s in area where we demand correctness and determinism that they will not be suitable. I think the thrust of this article is hard to see unless you have some experience with formal methods and verification. Or else accept the authors’ explanations as truth.
- n144q 2y agoI'll believe that when ChatGPT stops making up APIs that have never ever existed in the history of a library. The dumbest intern doesn't do that. Which is the entire point of the article that your comment fail to address.
- ojhughes 2y agoIn fairness, I’ve never experienced this using Claude with Cursor.
- n144q 2y ago*yet.
- jondwillis 2y agoUse Cursor or something similar and feed it documentation as context. Problem solved.
- n144q 2y agoSure, I'll save some time by referring to the doc while writing the code myself the old-fashioned way.
- zwnow 2y agoLmao, you are the type of person actually believing these silicon valley bs. o3 is far, far away from AGI.
- aaroninsf 2y agoIndeed, this put me immediately in mind of Ximm's Law: Every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon. Lemma: any statement about AI which uses the word "never" to preclude some feature from future realization is false.
- deleted 2y ago[deleted]
- layer8 2y agoAnd every advocate of AI assumes that it will necessarily and reasonably swiftly be improved to the level of AGI. Maybe assume neither?
- cootsnuck 2y agoBut o3 is just a slightly less stupid idiot savant...it still has to brute force solutions. Don't get me wrong, it's cool to see how far that technique can get you on a specific benchmark. But the point still stands that these systems can't be treated as deterministic (i.e. reliable or trustworthy) for the purposes of carrying out tasks that you can't allow "brute forced attempts" for (e.g. anything where the desired outcome is a positive subjective experience for a human). A new architecture is going to be needed that actually does something closer to our inherently heuristic based learning and reasoning. We'll still have the stochastic problem but we'll be moving further away from the idiot savant problem. All of this being said, I think there's plenty of usefulness with current LLMs. We're just expecting the wrong things from them and therefore creating suboptimal solutions. (Not everyone is, but the most common solutions are, IMO.) The best solutions need to be rethinking how we typically use software since software has been hinged upon being able to expect (and therefore test) dertiministic outputs from a limited set of user inputs. I work for an AI company that's been around for a minute (make our own models and everything). I think we're both in an AI hype bubble while simultaneously underestimating the benefits of current AI capabilities. I think the most interesting and potentially useful solutions are inherently going to be so domain specific that we're all still too new at realizing we need to reimagine how to build with this new tech in mind. It reminds me of the beginning of mobile apps. It took awhile for most us to "get it".
- turboat 2y agoCan you elaborate about your predictions for how the benefits of current capabilities will be applied? And your thoughts on how to build with it?
- JohnMakin 2y ago> After reading about o3's performance on ARC-AGI, I strongly suspect people will not be so flippantly dismissive of the inherent limits of these technologies by this time next year. If I wasn't so slammed with work I have half a mind to go dredge up at least a dozen posts that said the same thing last year, and the year before. Even OpenAI has been moving the goalposts here.
- tsurba 2y agoMy favorite quote in this topic: ”If intelligence lies in the process of acquiring new skills, there is no task X that solving X proves intelligence” IMO it especially applies to things like solving a new IQ puzzle, especially when the model is pretrained for that particular task type, like was done with ARC-AGI. For sure, it’s very good research to figure out what kind of tasks are easy for humans and difficult for ML, and then solve them. The jump in accuracy was surprising. But still in practice the models are unbeliavably stupid and lacking in common sense. My personal (moving) goalpost for ”AGI” is now set to whether a robot can keep my house clean automatically. Its not general intelligence if it can’t do the dishes. And before physical robots, being less of a turd at making working code would be a nice start. I’m not yet convinced general purpose LLMs will lead to cost-effective solutions to either vs humans. A specifically built dish washer however…
- cormackcorn 2y agoMyopic? You must be under 20 years old. For those of us who have been in tech for over four decades the OPs assessment is exactly the right framing.
- benterix 2y ago> After reading about o3's performance I heard that people still believing in OpenAI hype exist but I haven't met any IRL.