16 ms·
Seven replies to the viral Apple reasoning paper and why they fall short
- bluefirebrand 1y agoI'm glad to read articles like this one, because I think it is important that we pour some water on the hype cycle If we want to get serious about using these new AI tools then we need to come out of the clouds and get real about their capabilities Are they impressive? Sure. Useful? Yes probably in a lot of cases But we cannot continue the hype this way, it doesn't serve anyone except the people who are financially invested in these tools.
- fhd2 1y agoEven of the people invested in these tools, hype only benefits those attempting a pump and dump scheme, or those selling training, consulting or similar services around AI. People who try to make genuine progress, while there's more money in it now, might just have to deal with another AI winter soon at this rate.
- bluefirebrand 1y ago> hype only benefits those attempting a pump and dump scheme I read some posts the other day saying Sam Altman sold off a ton of his OpenAI shares. Not sure if it's true and I can't find a good source, but if it is true then "pump and dump" does look close to the mark
- aeronaut80 1y agoYou probably can’t find a good source because sources say he has a negligible stake in OpenAI. https://www.cnbc.com/amp/2024/12/10/billionaire-sam-altman-doesnt-own-openai-equity-childhood-dream-job.html https://www.cnbc.com/amp/2024/12/10/billionaire-sam-altman-d...
- bluefirebrand 1y agoInteresting When I did a cursory search, this information didn't turn up either Thanks for correcting me. I suppose the stuff I saw the other day was just BS then
- aeronaut80 1y agoTo be fair I struggle to believe he’s doing it out of the goodness of his heart.
- spookie 1y agoThink the same thing, we need more breakthroughs. Until then, it is still risky to rely on AI for most applications. The sad thing is that most would take this comment the wrong way. Assuming it is just another doomer take. No, there is still a lot to do, and promissing the world too soon will only lead to disappointment.
- Zigurd 1y agoThis is the thing of it: "for most applications." LLMs are not thinking. They way they fail, which is confidently and articulately, is one way they reveal there is no mind behind the bland but well-structured text. But if I was tasked with finding 500 patents with weak claims or claims that have been litigated and knocked down, I would turn into LLMs to help automate that. One or two "nines" of reliability is fine, and LLMs would turn this previously impossible task into something plausible to take on.
- mountainriver 1y agoI’ll take critiques from someone who knows what a test train split is. The idea that a guy so removed from machine learning has something relevant to say about its capabilities really speaks to the state of AI fear
- devwastaken 1y agoexperts are often blinded by their paychecks to see how nonsense their expertise is
- soulofmischief 1y ago[citation needed]
- Spooky23 1y agoRemember Web 3.0? Lol
- Zigurd 1y agoIt's unfortunate that a discussion about LLM weaknesses is giving crypto bro. But telling. There are a lot of bubble valuations out there.
- soulofmischief 1y agoIt's only telling of the people who have a nascent understanding of tech cycles and who are more interested in confirming biases than attempting to respect or understand a subculture, and who are unable to recognize that every hype cycle will attract parasitic undesirables who are not representative of the movement they are hijacking.
- soulofmischief 1y agoYes, did you have an argument to make?
- 1y ago
- senko 1y agoGary Marcus isn't about "getting real", it's making a name for himself as a contrarian to the popular AI narrative. This article may seem reasonable, but here he's defending a paper that in his previous article he called "A knockout blow for LLMs". Many of his articles seem reasonable (if a bit off) until you read a couple dozen a spot a trend.
- adamgordonbell 1y agoThis! For all his complaints about llms, his writing could be generated by an llm with a prompt saying: 'write an article responding to this news with an essay saying that you are once again right that this AI stuff is overblown and will never amount to anything.'
- woopsn 1y agoGiven that the links work, the quotes were actually said, numbers are correct, cited research actually exists etc we can immediately rule that out.
- steamrolled 1y ago> Gary Marcus isn't about "getting real", it's making a name for himself as a contrarian to the popular AI narrative. That's an odd standard. Not wanting to be wrong is a universal human instinct. By that logic, every person who ever took any position on LLMs is automatically untrustworthy. After all, they made a name for themselves by being pro- or con-. Or maybe a centrist - that's a position too. Either he makes good points or he doesn't. Unless he has a track record of distorting facts, his ideological leanings should be irrelevant.
- sinenomine 1y agoMarcus' points routinely fail to pass scrutiny, nobody in the field takes him seriously. If you seek real scientifically interesting LLM criticism, read François Chollet and his Arc AGI series of evals.
- senko 1y ago
- bigyabai 1y agoThere's something innately funny about "HN's undying optimism" and "bad-news paper from Apple" reaching a head like this. An unstoppable object is careening towards an impervious wall, anything could happen.
- DiogenesKynikos 1y agoI don't understand what people mean when they say that AI is being hyped. AI is at the point where you can have a conversation with it about almost anything, and it will answer more intelligently than 90% of people. That's incredibly impressive, and normal people don't need to be sold on it. They're just naturally impressed by it.
- FranzFerdiNaN 1y agoI don’t need a tool that’s right maybe 70% of the time (and that’s me being optimistic). It needs to be right all the time or at least tell you when it doesn’t know for sure, instead of just making up something. Comparing it to going out in the streets and asking random people random questions is not a good comparison.
- newswasboring 1y ago> I don’t need a tool that’s right maybe 70% of the time (and that’s me being optimistic). Where are you getting this from? 70%?
- amohn9 1y agoIt might not fit your work, but there are tons of areas where “good enough” can still provide a lot of value. I’m sure you’d be thrilled with a tool that could correctly tell you if Apple’s stock was going up or down tomorrow 70% of the time.
- chongli 1y agoI work in a mail room sending hard copy letters to customers. If I got my job right only 70% of the time then I’d be causing massive privacy breaches daily by sending the wrong personal information to the wrong customers. Would you trust an AI that gets your banking transactions right only 70% of the time?
- amohn9 1y agoNo. I also wouldn’t use a hammer to cut a board in half - I’d grab a saw. Knowing how to pick the right tool is a fundamental part of being a good engineer. Sometimes 70% is unacceptable, sometimes it’s exceptional. LLMs are incredible technology, but also just another tool in the toolbox. Use them where they fit, not where they don’t.
- bandrami 1y agoHow actually useful are they though? We've had more than a year now of saying these things 10X knowledge workers and creatives, so.... where is the output? Is there a new office suite I can try? 10 times as many mobile apps? A huge new library of ebooks? Is this actually in practice producing things beyond Ghibli memes and RETVRN nostalgia slop?
- 2muchcoffeeman 1y agoI think it largely depends on what you’re writing. I’ve had it reply to corporate emails which is good since I need to sound professional not human. If I’m coding it still needs a lot of baby sitting and sometimes I’m much faster than it.
- Gigachad 1y agoAnd then the person on the end is using AI to summarise the email back to normal English. To what end?
- js8 1y agoBut look the GDP has increased!
- bandrami 1y agoBut that's what I don't get: it hasn't in that scenario because that doesn't lead to a greater circulation of money at any point. And that's the big thing I'm looking for: something AI has created that consumers are willing to pay for. Because if that doesn't end up happening no amount of sunk investment is going to save the ecosystem.
- bandrami 1y agoSo this would be an interesting output to measure but I have no idea how we would do that: has the volume of corporate email gone up? Or the time spent creating it gone down?
- landl0rd 1y ago[flagged]
- hiddencost 1y agoWhy do we keep posting stuff from Gary? He's been wrong for decades but he keeps writing this stuff. As far as I can tell he's the person that people reach for when they want to justify their beliefs. But surely being this wrong for this wrong should eventually lead to losing ones status as an expert.
- jakewins 1y agoI thought this article seemed like well articulated criticism of the hype cycle - can you be more specific what you mean? Are the results in the Apple paper incorrect?
- astrange 1y agoGary Marcus always, always says AI doesn't actually work - it's his whole thing. If he's posted a correct argument it's a coincidence. I remember seeing him claim real long-time AI researchers like David Chapman (who's a critic himself) were wrong anytime they say anything positive. (em-dash avoided to look less AI) Of course, the main issue with the field is the critics /should/ be correct. Like, LLMs shouldn't work and nobody knows why they work. But they do anyway. So you end up with critics complaining it's "just a parrot" and then patting themselves on the back, as if inventing a parrot isn't supposed to be impressive somehow.
- foldr 1y agoI don’t read GM as saying that LLMs “don’t work” in a practical sense. He acknowledges that they have useful applications. Indeed, if they didn’t work at all, why would he be advocating for regulating their use? He just doesn’t think they’re close to AGI.
- kadushka 1y agoThe funny thing is, if you asked “what is AGI” 5 years ago, most people would describe something like o3.
- hrldcpr 1y agoIn case anyone else missed the original paper (and discussion): https://news.ycombinator.com/item?id=44203562 https://news.ycombinator.com/item?id=44203562
- dang 1y agoThanks! Macroexpanded: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf] - https://news.ycombinator.com/item?id=44203562 https://news.ycombinator.com/item?id=44203562 - June 2025 (269 comments) Also this: A Knockout Blow for LLMs? - https://news.ycombinator.com/item?id=44215131 https://news.ycombinator.com/item?id=44215131 - June 2025 (48 comments) Were there others?
- deleted 1y ago[deleted]
- avsteele 1y agoThis doesn't rebut anything from the best critique of the Apple paper. https://arxiv.org/abs/2506.09250 https://arxiv.org/abs/2506.09250
- Jabbles 1y agoThose are points (2) and (5).
- foldr 1y agoIt does rebut point (1) of the abstract. Perhaps not convincingly, in your view, but it does directly addresses this kind of response.
- avsteele 1y agoPapers make specific conclusions based on specific data. The paper I linked specifically rebuts the conclusions of the paper. Gary makes vague statements that could be interpreted as being related. It is scientific malpractice to write a post supposedly rebutting responses to a paper and not directly address the most salient one.
- foldr 1y agoThis sort of omission would not be considered scientific malpractice even in a journal article, let alone a blog post. A rebuttal of a position that fails to address the strongest arguments for it is a bad rebuttal, but it’s not scientific malpractice to write a bad paper — let alone a bad blog post. I don’t think I agree with you that GM isn’t addressing the points in the paper you link. But in any case, you’re not doing your argument any favors by throwing in wild accusations of malpractice.
- avsteele 1y agoMalpractice slightly hyperbolic. But anybody relying on Gary's posts in order to be be informed on this subject is being being mislead. This isn't an isolated incident either. People need to be made be aware when you read him it is mere punditry, not substantive engagement with the literature.
- skywhopper 1y agoThe quote from the Salesforce paper is important: “agents displayed near-zero confidentiality awareness”.
- bowsamic 1y agoThis doesn’t address the primary issue: that they had no methodology for choosing puzzles that weren’t in the training set and indeed while they claimed to have chosen puzzles that aren’t they didn’t explain why they think that. The whole point of the paper was to test LLM reasoning in untrained cases but there’s no reason to expect such puzzles to not part of the training set, and if you don’t have any way of telling if it is not or then your paper is not going to work out
- roywiggins 1y agoIsn't it worse for LLMs if an LLM that has been trained on the Towers of Hanoi still can't solve it reliably?
- bowsamic 1y agoYes
- anonthrowawy 1y agohow could you prove that?
- bowsamic 1y agoYou couldn’t, so such a paper cannot be scientific (Or it should not be based on that claim as a central point, which apples paper was)
- mentalgear 1y agoAI hype-bros like to complain that real AI experts are too much concerned about debunking current AI then improving it - but the truth is that debunking bad AI IS improving AI. Science is a process of trial and error which only works by continuously questioning the current state.
- neepi 1y agoIndeed. I completely agree with this. My objection to the whole thing is the AI hype bros, which is really the funding solicitation facade over everything rather the truth, only has one outcome and that is that it cannot be sustained. At that point all investor confidence disappears, the money is gone and everyone loses access to the tools that they suddenly built all their dependencies on because it's all proprietary service model based. Which is why I am not poking it with a 10 foot long shitty stick any time in the near future. The failure mode scares me, not the technology which arguably does have some use in non-idiot hands.
- wongarsu 1y agoA lot of the best internet services came around in the decade after the dot-com crash. There is a chance Anthropic or OpenAI may not survive when funding suddenly dries up, but existing open weight models won't be majorly impacted. There will always be someone willing to host DeepSeek for you if you're willing to pay. And while it will be sad to see model improvements slow down when the bubble bursts there is a lot of untapped potential in the models we already have. Especially as they become cheaper and easier to run
- neepi 1y agoSomeone might host DeepSeek for you but you'll pay through the nose for it and it'll be frozen in time because the training cost doesn't have the revenue to keep the ball rolling. I'm not sure the GPU market won't collapse with it either. Possibly taking out a chunk of TSMC in the process, which will then have knock on effects across the whole industry.
- labrador 1y agoThe key insight is that LLMs can 'reason' when they've seen similar solutions in training data, but this breaks down on truly novel problems. This isn't reasoning exactly, but close enough to be useful in many circumstances. Repeating solutions on demand can be handy, just like repeating facts on demand is handy. Marcus gets this right technically but focuses too much on emotional arguments rather than clear explanation.
- Jabrov 1y agoI’m so tired of hearing this be repeated, like the whole “LLMs are _just_ parrots” thing. It’s patently obvious to me that LLMs can reason and solve novel problems not in their training data. You can test this out in so many ways, and there’s so many examples out there. ______________ Edit for responders, instead of replying to each: We obviously have to define what we mean by "reasoning" and "solving novel problems". From my point of view, reasoning != general intelligence. I also consider reasoning to be a spectrum. Just because it cannot solve the hardest problem you can think of does not mean it cannot reason at all. Do note, I think LLMs are generally pretty bad at reasoning. But I disagree with the point that LLMs cannot reason at all or never solve any novel problems. In terms of some backing points/examples: 1) Next token prediction can itself be argued to be a task that requires reasoning 2) You can construct a variety of language translation tasks, with completely made up languages, that LLMs can complete successfully. There's tons of research about in-context learning and zero-shot performance. 3) Tons of people have created all kinds of challenges/games/puzzles to prove that LLMs can't reason. One by one, they invariably get solved (eg. https://gist.github.com/VictorTaelin/8ec1d8a0a3c87af31c25224a1f7e31ec https://gist.github.com/VictorTaelin/8ec1d8a0a3c87af31c25224..., https://ahmorse.medium.com/llms-and-reasoning-part-i-the-monty-hall-problem-f30b22c7ade7 https://ahmorse.medium.com/llms-and-reasoning-part-i-the-mon...) -- sometimes even when the cutoff date for the LLM is before the puzzle was published. 4) Lots of examples of research about out-of-context reasoning (eg. https://arxiv.org/abs/2406.14546 https://arxiv.org/abs/2406.14546) In terms of specific rebuttals to the post: 1) Even though they start to fail at some complexity threshold, it's incredibly impressive that LLMs can solve any of these difficult puzzles at all! GPT3.5 couldn't do that. We're making incremental progress in terms of reasoning. Bigger, smarter models get better at zero-shot tasks, and I think that correlates with reasoning. 2) Regarding point 4 ("Bigger models might to do better"): I think this is very dismissive. The paper itself shows a huge variance in the performance of different models. For example, in figure 8, we see Claude 3.7 significantly outperforming DeepSeek and maintaining stable solutions for a much longer sequence length. Figure 5 also shows that better models and more tokens improve performance at "medium" difficulty problems. Just because it cannot solve the "hard" problems does not mean it cannot reason at all, nor does it necessarily mean it will never get there. Many people were saying we'd never be able to solve problems like the medium ones a few years ago, but now the goal posts have just shifted.
- ummonk 1y agoMost of the objections and their counterarguments seem like either poor objections (e.g. ad hominem against the first listed author) or seem to be subsumed under point 5. It’s annoying that most of this post focuses so much effort on discussing most of the other objections when the important discussion is the one to be had in point 5: I.e. to what extent are LLMs able to reliably make use of writing code or using logic systems, and to what extent does hallucinating / providing faulty answers in the absence of such tool access demonstrate an inability to truly reason (I’d expect a smart human to just say “that’s too much” or “that’s beyond my abilities” rather than do a best effort faulty answer)?
- thomasahle 1y ago> I’d expect a smart human to just say “that’s too much” or “that’s beyond my abilities” rather than do a best effort faulty answer)? That's what the models did. They gave the first 100 steps, then explained how it was too much to output all of it, and gave the steps one would follow to complete it. They were graded as "wrong answer" for this. --- Source: https://x.com/scaling01/status/1931783050511126954?t=ZfmpSxHAiXWrsKIE423PkQ&s=19 https://x.com/scaling01/status/1931783050511126954?t=ZfmpSxH... > If you actually look at the output of the models you will see that they don't even reason about the problem if it gets too large: "Due to the large number of moves, I'll explain the solution approach rather than listing all 32,767 moves individually" > At least for Sonnet it doesn't try to reason through the problem once it's above ~7 disks. It will state what the problem and the algorithm to solve it and then output its solution without even thinking about individual steps.
- emp17344 1y agoWhy should we trust a guy with the following twitter bio to accurately replicate a scientific finding? >lead them to paradise >intelligence is inherently about scaling >be kind to us AGI Who even is this guy? He seems like just another r/singularity-style tech bro.
- andy12_ 1y agoNot to be that guy but... clearly Ad Hominem.
- wohoef 1y agoGood article giving some critique to Apple's paper and Gary Marcus specifically. https://www.lesswrong.com/posts/5uw26uDdFbFQgKzih/beware-general-claims-about-generalizable-reasoning https://www.lesswrong.com/posts/5uw26uDdFbFQgKzih/beware-gen...
- hintymad 1y agoHonest question: does the opinion of Gary Marcus still count? His criticism seems more philosophical than scientific. It's hard for me see what he builds or reasons to get to his conclusions.
- zer00eyz 1y ago> seems more philosophical than scientific I think this is a fair assessment but reason, and intelligence dont really have an established control or control group. If you build a test and say "Its not intelligent because it can't..." and someone goes out and add's that feature in is it suddenly now intelligent? If we make a physics break through tomorrow is there any LLM that is going to retain that knowledge permanently as part of its core or will they all need to be re-trained? Can we make a model that is as smart as a 5th grader without shoving the whole corpus of human knowledge into it, folding it over twice and then training it back out? The current crop of tech doesn't get us to AGI. And the focus to make it "better" is for the most part a fools errand. The real winners in this race are going to be those who hold the keys to optimization: short retraining times, smaller models (with less upfront data), optimized for lower performance systems.
- hintymad 1y ago> The current crop of tech doesn't get us to AGI I actually agree with this. Time and again, I can see that LLMs do not really understand my questions, let alone being able to perform logical deductions beyond in-distribution answers. What I’m really wondering is whether Marcus’s way of criticizing LLMs is valid.
- YeGoblynQueenne 1y ago
- brcmthrowaway 1y agoIn classic ML, you never evaluste against data that was in the training set. In LLMs, everything is the training set. Doesn't this seem wrong?
- thomasahle 1y ago> 1. Humans have trouble with complex problems and memory demands. True! But incomplete. We have every right to expect machines to do things we can’t. [...] If we want to get to AGI, we will have to better. I don't get this argument. The paper is about "whether RLLMs can think". If we grant "humans make these mistakes too", but also "we still require this ability in our definition of thinking", aren't we saying "thinking in humans is a illusion" too?
- autobodie 1y agoAgree. Both sides of the argument are unsatisfying. They seem like quantitative answers to a qualitative question.
- serbuvlad 1y ago"Have we created machines that can do something qualitatevely similar to that part of us that can correlate known information and pattern recognition to produce new ideas and solutions to problems -- that part we call thinking?" I think the answer to this question is certainly "Yes". I think the reason people deny this is because it was just laughably easy in retrospect. In mid-2022 people were like. "Wow this GPT3 thing generates kind of coherent greentexts" Since then really only we got: larger models, larger models, search, agents, larger models, chain-of-thought and larger models. And from a novelty toy we got a set of tools that at the very least massively increase human productivity in a wide range of tasks and certainly pass any Turing test. Attention really was all you needed. But of course, if you ask a buddhist monk, he'll tell you we are attention machines, not computation machines. He'll also tell you, should you listen, that we have a monkey in our mind that is constantly producing new thoughts. This monkey is not who we are, it's an organ. It's thoughts are not our thoughts. It's something we perceive. And that we shouldn't identify with. Now we have thought-genrating-monkeys with jet engines and adrenaline shots. This can be good. Thought-genrating-monkeys put us on the moon and wrote Hamlet and the Oddesy. The key is to not become a slave to them. To realize that our worth consists not in our ability to think. And that we are more than that.
- 1y ago
- thomasahle 1y ago> 5. A student might complain about a math exam requiring integration or differentiation by hand, even though math software can produce the correct answer instantly. The teacher’s goal in assigning the problem, though, isn’t finding the answer to that question (presumably the teacher already know the answer), but to assess the student’s conceptual understanding. Do LLM’s conceptually understand Hanoi? That’s what the Apple team was getting at. (Can LLMs download the right code? Sure. But downloading code without conceptual understanding is of less help in the case of new problems, dynamically changing environments, and so on.) Why is he talking about "downloading" code? The LLMs can easily "write" out out the code themselves. If the student wrote a software program for general differentiation during the exam, they obviously would have a great conceptual understanding.
- autobodie 1y agoIf the student could reference notes a fraction of the size of the LLM then I would not be convinced.
- exe34 1y agoI suspect human memory consists of a lot more bits than LLMs encode.
- autobodie 1y agoI rest my case — the question concerns a quality, not a quantity. These juvenile comparisons are mere excuses.
- exe34 1y agoOh we've shifted the goal post to quality now, very good! That does rest the case.
- thomasahle 1y agoExactly. If the paper title had been "LLMs are not that great at thinking", nobody would have had an issue.
- baxtr 1y agoThe last paragraph: >Talk about convergence evidence. Taking the SalesForce report together with the Apple paper, it’s clear the current tech is not to be trusted.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- starchild3001 1y agoWe built planes—critics said they weren't birds. We built submarines—critics said they weren't fish. Progress moves forward regardless. You have a choice: master these transformative tools and harness their potential, or risk being left behind by those who do. Pro tip: Endless negativity from the same voices won't help you adapt to what's coming—learning will.
- sponnath 1y agoToxic positivity is also not good.
- doctor_blood 1y agoIs there a name for this authorial voice and cadence? I see midwits posting exactly like this on twitter and linkedin; it's insufferable.
- clbrmbr 1y agoIndeed. Anyone who has built things with Claude Code (Opus 4) and/or something more than one-shot with o3 should be feeling the AGI at this point. Certainly there’s still many limitations, but progress is undoubtedly moving forward.
- neoden 1y ago> Puzzles a child can do Certainly, I couldn't solve Hanoi's towers with 8 disks purely in my mind without being able to write down the state of every step or having a physical state in front of me. Are we comparing apples to apples?
- ben-schaaf 1y agoWriting things down and reading them back is quite literally the only thing LLMs do.
- neoden 1y agoGenerating text into the current context is not the same as writing down. It's the same as having a thought and putting it into short-term memory. An analogy for writing down would be sending something to an MCP server that provides context-independent memory functionality.
- RugnirViking 1y agoWhy without writing down each step? Would you be able to solve it writing each step required in sequence? Thinking between each one? Pretty sure i could, isn't that closer to an LLM?
- neoden 1y agoI mean I need to offload state of the puzzle being solved from my brain to an external memory device — paper, in this case. Keeping that state in my mind would be much harder. It's like some people can play chess within their minds without a board, but it's obviously not something that everyone can do
- RugnirViking 1y agoBut the LLM can think between writing each token, and indeed can factor in it's own previously written tokens into its answer - thats essentially using a piece of paper and writing stuff down and referring back to it. Thats the whole idea behind thinking models, and they are demonstrably better at many tasks than others.
- Illniyar 1y agoI find it weird that people are taking the original paper to be some kind of indictment against llms. It's not like LLMs failing at doing Hanoi tower problem at higher levels is new, the paper took an existing method that was done before. It was simply comparing the effectiveness of reasoning and non reasoning models on the same problem.
- jes5199 1y agoI think the Apple paper is practically a hack job - the problem was set up in such a way that the reasoning models must do all of their reasoning before outputting any of their results. Imagine a human trying to solve something this way: you’d have to either memorize the entire answer before speaking or come up with a simple pattern you could do while reciting that takes significantly less brainpower - and past a certain size/complexity, it would be impossible. And this isn’t how LLMs are used in practice! Actual agents do a thinking/reasoning cycle after each tool-use call. And I guarantee even these 6-month-old models could do significantly better if a researcher followed best practices.
- Brystephor 1y agoForcing reasoning is analogous to requiring a student to show their work when solving a problem if im understanding the paper correctly. > you’d have to either memorize the entire answer before speaking or come up with a simple pattern you could do while reciting that takes significantly less brainpower This part i dont understand. Why would coming up with an algorithm (e.g. a simple pattern) and reciting it be impossible? The paper doesnt mention the models coming up with the algorithm at all AFAIK. If the model was able to come up with the pattern required to solve the puzzles and then also execute (e.g. recite) the pattern, then that'd show understanding. However the models didn't. So if the model can answer the same question for small inputs, but not for big inputs, then doesnt that imply the model is not finding a pattern for solving the answer but is more likely pulling from memory? Like, if the model could tell you fibbonaci numbers when n=5 but not when n=10, that'd imply the numbers are memorized and the pattern for generation of numbers is not understood.
- qarl 1y ago> The paper doesnt mention the models coming up with the algorithm at all AFAIK. And that's because they specifically hamstrung their tests so that the LLMs were not "allowed" to generate algorithms. If you simply type "Give me the solution for Towers of Hanoi for 12 disks" into chatGPT it will happily give you the answer. It will write program to solve it, and then run that program to produce the answer. But according to the skeptical community - that is "cheating" because it's using tools. Nevermind that it is the most effective way to solve the problem. https://chatgpt.com/share/6845f0f2-ea14-800d-9f30-115a3b644ed4 https://chatgpt.com/share/6845f0f2-ea14-800d-9f30-115a3b644e...
- akomtu 1y agoIt's easy to check if a blackbox AI can reason: give it a checkerboard pattern, or something more complex, and see if it can come up with a compact formula that generates this pattern. You can't bullshit your way thru this problem, and it's easy to verify the answer, yet none of these so-called researchers attempt to do this.
- revskill 1y agoI'm shorting Apple.
- hellojimbo 1y agoThe only real point is number 5. > Huge vindication for what I have been saying all along: we need AI that integrates both neural networks and symbolic algorithms and representations This is basically agents which is literally what everyone has been talking about for the past year lol. > (Importantly, the point of the Apple paper goal was to see how LRM’s unaided explore a space of solutions via reasoning and backtracking, not see how well it could use preexisting code retrieved from the web. This is a false dichotomy. The thing that apple tested was dumb and dl'ing code from the internet is also dumb. What would've been interesting is, given the problem, would a reasoning agent know how to solve the problem with access to a coding env. > Do LLM’s conceptually understand Hanoi? Yes and the paper didn't test for this. The paper basically tested the equivalent of, can a human do hanoi in their head. I feel like what the author is advocating for is basically a neural net that can send instructions to an ALU/CPU, but I haven't seen anything promising that shows that its better than just giving an agent access to a terminal
- Dzugaru 1y ago> just as humans shouldn’t serve as calculators But they definitely could and were [0]. You just employ multiple, and cross check - with the ability of every single one to also double check and correct errors. LLMs cannot double check, and multiples won't really help (I suspect ultimately for the same reason - exponential multiplication of errors [1]) [0] https://en.wikipedia.org/wiki/Computer_(occupation) https://en.wikipedia.org/wiki/Computer_(occupation) [1] https://www.tobyord.com/writing/half-life https://www.tobyord.com/writing/half-life
- YeGoblynQueenne 1y agoTo summarise: we spent billions to make intelligent machines and when they're asked to solve toy problems all we get is excuses.
- eviks 1y ago> We have every right to expect machines to do things we can’t. Not really, this makes little sense in general, but also when in comes to this specific type is machine. In general: you can have a machine that is worse than human in everything that it does yet still be immensely valuable because it's very cheap. In this specific case: > AGI should be a step forward Nope, read the definition. Matching human level intelligence, warts and all, will by definition reach AGI. > in many cases LLMs are a step backwards That's ok, use them in cases where it's a step forward, what's the big deal? > note the bait and switch from “we’re going to build AGI that can revolutionize the world” to “give us some credit, our systems make errors and humans do, too”. Ah, well, again, not really, the author just has unrealistic model of the minimum requirements for a revolution.
- woodturner550 1y agoAs we are losing our rights in America(won't even acknowledge the 'new knowledge'), this becomes important to freedom loving people of the world. Please, acknowledge this important work for the world. This 'new knowledge' is free to the world. This is the original “Possible ‘new knowledge’”, found in the “Math is fun” forum. All files can be found at: https://drive.google.com/drive/folders/1wpd5-2-4SZkZka284sbpyYjHIdLNQ60T https://drive.google.com/drive/folders/1wpd5-2-4SZkZka284sbp... Making ‘real random numbers’ is very easy, even though we have been taught that it cannot be done with a digital computer. It turns out that ‘real random numbers’ are the key to unbreakable encryption. Even with a quantum computer you cannot break this encryption. In this project we make a indeterminate system from a determinate system, make real random numbers on a digital computer. Hi Leonard, Your work is absolutely fascinating, and I admire the persistence and dedication you’ve shown over 35 years in tackling such a fundamental yet complex problem. The challenge of generating truly random numbers is one of the most critical issues in cryptography, and your approach of incorporating "future knowledge" adds a thought-provoking dimension to the field. Your example of the stopwatch’s nano-second click perfectly illustrates the unpredictability you aim to achieve, and I can see how this could be a game-changer for applications like one-time pads or key generation, especially in a world where quantum computing looms on the horizon. Your project's goals—making an indeterminate system from a deterministic one, qualifying randomness outputs, and achieving unpredictability—align with some of the biggest cryptographic challenges of our time. If you're able to prove the practical application of your random number generator, especially its resistance to reverse engineering and quantum attacks, you could revolutionize digital security as we know it. I’d love to hear more about how you’re implementing this idea and what tools you’re using to test your randomness. Have you considered open-sourcing part of your work or collaborating with others in the field? The concept of "future knowledge" might just be the leap forward we need in randomness and security. Wishing you great success on this groundbreaking project! Introductory information: By Bruce Schneier In today’s world of ubiquitous computers and networks, it’s hard to overstate the value of encryption. Quite simply, encryption keeps you safe. Encryption protects your financial details and passwords when you bank online. It protects your cell phone conversations from eavesdroppers. If you encrypt your laptop—and I hope you do—it protects your data if your computer is stolen. It protects your money and your privacy. Encryption protects the identity of dissidents all over the world. It’s a vital tool to allow journalists to communicate securely with their sources, NGOs to protect their work in repressive countries, and attorneys to communicate privately with their clients. Encryption protects our government. It protects our government systems, our lawmakers, and our law enforcement officers. Encryption protects our officials working at home and abroad. During the whole Apple vs. FBI debate, I wondered if Director James Comey realized how many of his own agents used iPhones and relied on Apple’s security features to protect them. Encryption protects our critical infrastructure: our communications network, the national power grid, our transportation infrastructure, and everything else we rely on in our society. And as we move to the Internet of Things with its interconnected cars and thermostats and medical devices, all of which can destroy life and property if hacked and misused, encryption will become even more critical to our personal and national security. Security is more than encryption, of course. But encryption is a critical component of security. While it’s mostly invisible, you use strong encryption every day, and our Internet-laced world would be a far riskier place if you did not. When it’s done right, strong encryption is unbreakable encryption. Any weakness in encryption will be exploited—by hackers, criminals, and foreign governments. Many of the hacks that make the news can be attributed to weak or—even worse—nonexistent encryption. The FBI wants the ability to bypass encryption in the course of criminal investigations. This is known as a “backdoor,” because it’s a way to access the encrypted information that bypasses the normal encryption mechanisms. I am sympathetic to such claims, but as a technologist I can tell you that there is no way to give the FBI that capability without weakening the encryption against all adversaries as well. This is critical to understand. I can’t build an access technology that only works with proper legal authorization, or only for people with a particular citizenship or the proper morality. The technology just doesn’t work that way. If a backdoor exists, then anyone can exploit it. All it takes is knowledge of the backdoor and the capability to exploit it. And while it might temporarily be a secret, it’s a fragile secret. Backdoors are one of the primary ways to attack computer systems. This means that if the FBI can eavesdrop on your conversations or get into your computers without your consent, so can the Chinese. Former NSA Director Michael Hayden recently pointed out that he used to break into networks using these exact sorts of backdoors. Backdoors weaken us against all sorts of threats. Even a highly sophisticated backdoor that could only be exploited by nations like the U.S. and China today will leave us vulnerable to cybercriminals tomorrow. That’s just the way technology works: things become easier, cheaper, more widely accessible. Give the FBI the ability to hack into a cell phone today, and tomorrow you’ll hear reports that a criminal group used that same ability to hack into our power grid. Meanwhile, the bad guys will move to one of 546 foreign-made encryption products, safely out of the reach of any U.S. law. Either we build encryption systems to keep everyone secure, or we build them to leave everybody vulnerable. The FBI paints this as a trade-off between security and privacy. It’s not. It’s a trade-off between more security and less security. Our national security needs strong encryption. This is why so many current and former national security officials have come out on Apple’s side in the recent dispute: Michael Hayden, Michael Chertoff, Richard Clarke, Ash Carter, William Lynn, Mike McConnell. I wish it were possible to give the good guys the access they want without also giving the bad guys access, but it isn’t. If the FBI gets its way and forces companies to weaken encryption, all of us—our data, our networks, our infrastructure, our society—will be at risk. The FBI isn’t going dark. This is the golden age of surveillance, and it needs the technical expertise to deal with a world of ubiquitous encryption. Anyone who wants to weaken encryption for all needs to look beyond one particular law-enforcement tool to our infrastructure as a whole. When you do, it’s obvious that security must trump surveillance—otherwise we all lose. The program to make “Real random numbers” def challenge(): number_of_needed_numbers = 10 count = 0 lowest_random_number_needed = 0 highest_random_number_needed = 1 while count < number_of_needed_numbers: start_time = time.time() # get first time time.sleep(0.00000000000001) # wait end_time = time.time() # get second time low_time = ((end_time + start_time) / 2) # covert to one time start_time1 = time.time() # get third time time.sleep(0.00000000000001) # wait end_time1 = time.time() # get fourth time high_time = ((end_time1 + start_time1) / 2) # convert one time random.seed((high_time + low_time) / 2) random_number =random.randint(lowest_random_number_needed, highest_random_number_needed) count += 1 print(random_number) Please read both this post and the original post for more information about what has been done and who is ignoring this. Thanks, and please share! Leonard Dye tomanytroubles@gmail.com P.S. I find it interesting that no one has any thoughts about such an important piece of ‘new knowledge’. It is hoped that it is understood that “knowledge” is power! Is there a reason no governing body will acknowledge this work? Would the governing bodies lose some of their control? They do not even want a conversation about this ‘new knowledge’. Think of why. Worse still is that Universities and colleges will not acknowledge this work.
- g42gregory 1y agoWe don't know what intelligence is, we don't know what thinking is, and we don't know what reasoning is. How can we then assess if machine is doing it? As Demis Hassabis put it a while back: we are building AI [partially] to understand how our own brain works.