6 ms·
I haven’t seen this new google model but now must try it out. I will say that other frontier models are starting to surprise me with their reasoning/understand
by efitz 11mo ago
I haven’t seen this new google model but now must try it out.
I will say that other frontier models are starting to surprise me with their reasoning/understanding- I really have a hard time making (or believing) the argument that they are just predicting the next word.
I’ve been using Claude Code heavily since April; Sonnet 4.5 frequently surprises me.
Two days ago I told the AI to read all the documentation from my 5 projects related to a tool I’m building, and create a wiki, focused on audience and task.
I'm hand reviewing the 50 wiki pages it created, but overall it did a great job.
I got frustrated about one issue: I have a github issue to create a way to integrate with issue trackers (like Jira), but it's TODO, and the AI featured on the home page that we had issue tracker integration. It created a page for it and everything; I figured it was hallucinating.
I went to edit the page and replace it with placeholder text and was shocked that the LLM had (unprompted) figured out how to use existing features to integrate with issue trackers, and wrote sample code for GitHub, Jira and Slack (notifications). That truly surprised me.
- energy123 11mo agoPredicting the next word requires understanding, they're not separate things. If you don't know what comes after the next word, then you don't know what the next word should be. So the task implicitly forces a more long-horizon understanding of the future sequence.
- IAmGraydon 11mo agoThis is utterly wrong. Predicting the next word requires a large sample of data made into a statistical model. It has nothing to do with "understanding", which implies it knows why rather than what.
- orionsbelt 11mo agoIlya Sustkever was on a podcast, saying to imagine a mystery novel where at the end it says “and the killer is: (name)”. Saying it’s just a statistical model generating the next most likely word, how can it do that in this case if it doesn’t have some understanding of all the clues, etc. A specific name is not statistically likely to appear
- shwaj 11mo agoCan current LLMs actually do that, though? What Ilya posed was a thought experiment: if it could do that, then we would say that it has understanding. But AFAIK that is beyond current capabilities.
- krackers 11mo agoSomeone should try it and create a new "mysterybench". Find all mystery novels written after LLM training cutoff, and see how many models unravel the mystery
- IAmGraydon 11mo agoIt can't do that without the answer to who did it being in the training data. I think the reason people keep falling for this illusion is that they can't really imagine how vast the training dataset is. In all cases where it appears to answer a question like the one you posed, it's regurgitating the answer from its training data in a way that creates an illusion of using logic to answer it.
- CamperBob2 11mo agoIt can't do that without the answer to who did it being in the training data. Try it. Write a simple original mystery story, and then ask a good model to solve it. This isn't your father's Chinese Room. It couldn't solve original brainteasers and puzzles if it were.
- dyauspitr 11mo agoThat’s not true, at all.
- IAmGraydon 11mo agoPlease…go on.
- stevenhuang 11mo agoYou sound more like a stochastic parrot than an LLM does at this point.
- astrange 11mo agoIf you're claiming a transformer model is a Markov chain, this is easily disprovable by, eg, asking the model why it isn't a Markov chain! But here is a really big one of those if you want it: https://arxiv.org/abs/2401.17377 https://arxiv.org/abs/2401.17377
- nl 11mo agoModern LLMs are post trained for tasks other than next word prediction. They still output words through (except for multi-modal LLMs) so that does involve next word generation.
- Workaccount2 11mo ago"Understanding" is just a trap to get wrapped up in. A word with no definition and no test to prove it. Whether or not the model are "understanding" is ultimately immaterial, as their ability to do things is all that matters.
- pinnochio 11mo agoIf they can't do things that require understanding, it's material, bub. And just because you have no understanding of what "understanding" means, doesn't mean nobody does.
- red75prime 11mo ago> doesn't mean nobody does If it's not a functional understating that allows to replicate functionality of understanding, is it the real understanding?
- dyauspitr 11mo agoThe line between understanding and “large sample of data made into a statistical model” is kind of fuzzy.
- HarHarVeryFunny 11mo ago> Predicting the next word requires understanding If we were talking about humans trying to predict next word, that would be true. There is no reason to suppose than an LLM is doing anything other than deep pattern prediction pursuant to, and no better than needed for, next word prediction.
- CamperBob2 11mo agoHow'd you do at the International Math Olympiad this year?
- HarHarVeryFunny 11mo agoI hear the LLM was able to parrot fragments of the stuff it was trained to memorize, and did very well
- CamperBob2 11mo agoYeah, that must be it.
- cxvrfr 11mo agoWell being able to extrapolate solutions to "novel" mathematical exercises based on a very large sample of similar tasks in your dataset seems like a reasonable explanation. Question is how well it would do if it was trained without those samples?
- CamperBob2 11mo agoGee, I don't know. How would you do at a math competition if you weren't trained with math books? Sample problems and solutions are not sufficient unless you can genuinely apply human-level inductive and deductive reasoning to them. If you don't understand that and agree with it, I don't see a way forward here. A more interesting question is, how would you do at a math competition if you were taught to read, then left alone in your room with a bunch of math books? You wouldn't get very far at a competition like IMO, calculator or no calculator, unless you happen to be some kind of prodigy at the level of von Neumann or Ramanujan.
- astrange 11mo agoPredicting the next word is the interface, not the implementation. (It's a pretty constraining interface though - the model outputs an entire distribution and then we instantly lose it by only choosing one token from it.)
- charcircuit 11mo agoIt's trying to maximize a reward function. It's not just predicting the next word.
- schiffern 11mo ago>I really have a hard time making (or believing) the argument that they are just predicting the next word. It's true, but by the same token our brain is "just" thresholding spike rates.