5 ms·
I believe Gary Marcus is quite well known for terrible AI predictions. He's not in any way an expert in the field. Some of his predictions from 2022 [1] > In 2
by gejose 9mo ago
I believe Gary Marcus is quite well known for terrible AI predictions. He's not in any way an expert in the field. Some of his predictions from 2022 [1]
> In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.
> In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.
> In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).
> In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
> In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
Many of these have already been achieved, and it's only early 2026.
[1]https://garymarcus.substack.com/p/dear-elon-musk-here-are-five-things https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi...
- ls612 9mo agoI'm pretty sure it can do all of those except for the one which requires a physical body (in the kitchen) and the one that humans can't do reliably either (construct 10000 loc bug-free).
- merlincorey 9mo agoWhich ones are you claiming have already been achieved? My understanding of the current scorecard is that he's still technically correct, though I agree with you there is velocity heading towards some of these things being proven wrong by 2029. For example, in the recent thread about LLMs and solving an Erdos problem I remember reading in the comments that it was confirmed there were multiple LLMs involved as well as an expert mathematician who was deciding what context to shuttle between them and helping formulate things. Similarly, I've not yet heard of any non-expert Software Engineers creating 10,000+ lines of non-glue code that is bug-free. Even expert Engineers at Cloud Flare failed to create a bug-free OAuth library with Claude at the helm because some things are just extremely difficult to create without bugs even with experts in the loop.
- stingrae 9mo ago1 and 2 have been achieved. 4 is close, the interface needs some work to allow nontechnical people use it. (claude code)
- fxtentacle 9mo agoI strongly disagree. I’ve yet to find an AI that can reliably summarise emails, let alone understand nuance or sarcasm. And I just asked ChatGPT 5.2 to describe an Instagram image. It didn’t even get the easily OCR-able text correct. Plus it completely failed to mention anything sports or stadium related. But it was looking at a cliche baseball photo taken by an fan inside the stadium.
- protocolture 9mo agoI have had ChatGPT read text in an image, give me a 100% accurate result, and then claim not to have the ability and to have guessed the previous result when I ask it to do it again.
- pixl97 9mo ago>let alone understand nuance or sarcasm I'm still trying to find humans that do this reliably too. To add on, 5.2 seems to be kind of lazy when reading text in images by default. Feeding it an image it may give the first word or so. But coming back with a prompt 'read all the text in the image' makes it do a better job. With one in particular that I tested I thought it was hallucinating some of the words, but there was a picture in the picture with small words it saw I missed the first time. I think a lot of AI capabilities are kind of munged to end users because they limit how much GPU is used.
- atomic_reed 9mo ago[dead]
- falloutx 9mo agoI dispute 1 & 2 more than 4. 1) Is it actually watching a movie frame by frame or just searching about it and then giving you the answer? 2) Again can it handle very long novels, context windows are limited and it can easily miss something. Where is the proof for this? 4 is probably solved 4) This is more on predictor because this is easy to game. you can create some gibberish code with LLM today that is 10k lines long without issues. Even a non-technical user can do
- zozbot234 9mo ago> In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI. Can AI actually do this? This looks like a nice benchmark for complex language processing, since a complete novel takes up a whole lot of context (consider War and Peace or The Count of Monte Cristo). Of course the movie variety is even more challenging since it involves especially complex multi-modal input. You could easily extend it to making sense of a whole TV series.
- the-grump 9mo agoYes they can. The size of many codebases is much larger and LLMs can handle those. Consider also that they can generate summaries and tackle the novel piecemeal, just like a human would. Re: movies. Get YouTube premium and ask YouTube to summarize a 2hr video for you.
- falloutx 9mo agoNovel is different from a codebase. In code you can have a relationship between files and most files can be ignored depending on what you're doing. But for a novel, its a sequential thing, in most cases A leads to B and B leads to C and so on. > Re: movies. Get YouTube premium and ask YouTube to summarize a 2hr video for you. This is different from watching a movie. Can it tell what suit actor was wearing? Can it tell what the actor's face looked like? Summarising and watching are too different things.
- cmcaleer 9mo agoYou’re moving the goalposts. Gary Marcus’ proposal was being able to ask: Who are the characters? What are their conflicts and motivations? etc. Which is a relatively trivial task for a current LLM.
- daveguy 9mo agoThe Gary Marcus proposal you refer to was about a novel, and not a codebase. I think GP's point is that motivations require analysis outside of the given (or derived) context window, which LLMs are essentially incapable of doing.
- colechristensen 9mo agoBesides being a cook which is more of a robotics problem all of the rest are accomplished to the point of being arguable about how reliably LLMs can perform these tasks, the arguments being between the enthusiast and naysayer camps. The keyword being "reliably" and what your threshold is for that. And what "bug free" means. Groups of expert humans struggle to write 10k lines of "bug free" code in the absolutist sense of perfection, even code with formal proofs can have "bugs" if you consider the specification not matching the actual needs of reality. All but the robotics one are demonstrable in 2026 at least.
- thethirdone 9mo agoWhich ones of those have been achieved in your opinion? I think the arbitrary proofs from mathematical literature is probably the most solved one. Research into IMO problems, and Lean formalization work have been pretty successful. Then, probably reading a novel and answering questions is the next most successful. Reliably constructing 10k bug free lines is probably the least successful. AI tends to produce more bugs than human programmers and I have yet to meet a programmer who can reliably produce less than 1 bug per 10k lines.
- zozbot234 9mo agoFormalizing an arbitrary proof is incredibly hard. For one thing, you need to make sure that you've got at least a correct formal statement for all the prereqs you're relying on, or the whole thing becomes pointless. Many areas of math ouside of the very "cleanest" fields (meaning e.g. algebra, logic, combinatorics etc.) have not seen much success in formalizing existing theory developments.
- kleene_op 9mo ago> Reliably constructing 10k bug free lines is probably the least successful. You imperatively need to try Claude Code, because it absolutely does that.
- thethirdone 9mo agoI have seen many people try to use Claude Code and get LOTS of bugs. Show me any > 10k project you have made with it and I will put the effort in to find one bug free of charge.
- jgalt212 9mo agoThis comment or something very close always appears alongside a Gary Marcus post.
- margalabargala 9mo agoWhich is fortunate, considering how asinine it is in 2026 to expect that none of the items listed will be accomplished in the next 3.9 years.
- GorbachevyChase 9mo agoI think it’s for good reason. I’m a bit at a loss as to why every time this guy rages into the ether of his blog it’s considered newsworthy. Celebrity driven tech news is just so tiresome. Marcus was surpassed by others in the field and now he’s basically a professional heckler on a university payroll. I wish people could just be happy for the success of others instead of fuming about how so and so is a billionaire and they are not.
- raincole 9mo agoAnd why not? Is there any reason for this comment to not appear? If Bill Gates made a predication about computing, no matter what the predication says, you can bet that 640K memory quote would be mentioned in the comment section (even he didn't actually say that).
- jgalt212 9mo agobecuase - it's tiresome - and the only less useful than making predictions is making predictions about predictions.
- raincole 9mo ago> Many of these have already been achieved, and it's only early 2026. I'm quite sure people who made those (now laughable) predictions will tell you none of these has been achieved, because AI isn't doing this "reliably" or "bug-free." Defending your predictions is like running an insurance company. You always win.
- dyauspitr 9mo agoIn my opinion, contrary to other comments here I think AI can do all of the above already except being a kitchen cook. Just earlier today I asked it to give me a summary of a show I was watching until a particular episode in a particular season without spoiling the rest of it and it did a great job.
- suddenlybananas 9mo agoYou know that almost every show as summaries of episodes available online?
- joquarky 9mo agoHow do you find them?
- staticman2 9mo agoI don't understand how this claim can even be tested: > In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI. Once you are "going beyond the literal text" the standard is usefulness of your insight about the novel, not whether your insight is "right" or "wrong".