7 ms·
LLMs struggle to explain themselves
- jonathanyc 2y agoHi HN! I'm about to go to bed, but I promise to go through comments when I wake up later today. Happy Friday! The source code for the demo is on GitHub: https://github.com/jyc/stackbee https://github.com/jyc/stackbee
- latexr 2y ago> One day in March, I was walking my dog, saw a house numbered 3147, and thought it was a funny pattern. It’s the Fibonacci sequence with a different seed (31 instead of 11). The Fibonacci sequence with 3,1 would be 3,1,4,5. I think you mean the house number was 1347. That would work and be easier to notice.
- jonathanyc 2y agoHaha thanks for catching that—fixed! That’s pretty ironic. Looks like I’m the LLM and you’re the human :P Now I can’t remember whether the house was 3147 or 1347. The pattern might have been to add the first digit to the last digit (unrot + in the stack language). That’s what I get for writing at 3am!
- beardyw 2y ago> It’s interesting to me that in spite of the fact that the LLM does so poorly with number sequences in general, it does pretty well with variations of the Fibonacci sequence. Not surprising, since the Fibonacci sequence will be in the text swallowed by the LLM.
- jonathanyc 2y agoYea! Do you think the LLM learns a Fibonacci-specific circuit? Do you think it is possible for an LLM to learn a general pattern-recognition circuit?
- beardyw 2y agoWell, a basic LLM would not really even distinguish a Fibonacci series from a hole in the ground. But models are being built with extra stuff, though I doubt recognising number series is high on the list.
- deleted 2y ago[deleted]
- reasonableklout 2y agoThanks for sharing! The custom stack-based language that you created for randomly generating interesting integer sequences was the most interesting part of this post for me. Wish the post focused on that rather than LLMs!
- EternalFury 2y agoEither you believe in emergent behavior or it’s only very good at recognizing patterns. Which is it? I made a test called LLMLOLinator (stupid name) and I can tell they cannot stray from the probability distributions they learned during training. So, I am not confident in so-called emergent behaviors.
- fsndz 2y agoThere are no emergent behaviors; LLMs are essentially memorizing statistical patterns in the data and using lexical cues to generate responses. They cannot explain themselves reliably because they don't know what they know, nor do they know what they don't know. In fact, LLMs don't truly 'know' anything at all. These are not thinking machines—they are simply the result of statistical pattern matching on steroids. That's why there will be no AGI, at least not from LLMs: https://www.lycee.ai/blog/why-no-agi-openai https://www.lycee.ai/blog/why-no-agi-openai
- d_sem 2y agoAnd yet I can provide modern LLMs unique never before asked scenarios that require the ability to reason about real world phenomenon and they can respond in ways more thoughtful than the average person. Much of human education is feeding books of information about things we will never experience in our day-to-day lives, and convincing ourselves it reflects reality. When in fact most of us have not personally experienced any evidences that what we learned is true. That vast majority of what a person "knows" is a biological statistical pattern matching on steroids.
- Salgat 2y agoWhat you're describing is no different than a linear regression describing the predicted value between two data points. Sometimes the regression is close if enough data exists, other times you get wild hallucinations, with the model none the wiser whether it's correct. All this tells us is that there is a lot of data out there on the internet that can still have useful information extracted from it. Take someone like Ramanujan, who with a couple math books on his own, could derive brilliant and novel discoveries in mathematics, instead of needing millions of man hours worth reading material to replicate what is mostly a replacement for googling.
- hun3 2y agoEven humans are bad at explaining their decisions sometimes, especially if they did not reach there by reason. In fact, if asked for the reason post mortem, people (and LLM) are likely to make them up on the spot. I wonder if the same dynamics is at play here.
- threeseed 2y agoLLMs will provide an answer for every question even when told it doesn't know what it is talking about. Overwhelming majority of people won't do this. So I don't think the comparison really makes any sense.
- adastra22 2y agoHave you used them recently? They have gotten much better about this.
- scratcheee 2y agoI’ve met humans who do that too
- jonathanyc 2y agoYea that’s a good point! On the one hand, it’s cool to look at the JavaScript the LLM generates to test its hypotheses in the logs (you can click “expand” to see them). On the other hand you’re right that especially given that it’s (presumably) an autoregressive LLM, if it writes the choice before the pattern there’s no way for the pattern to influence the choice. You can muck with the prompts by clicking “Settings”! I think there’s a lot of room for improvement with all aspects of the experiment. One prompt I tried made the LLM take tons of turns, but it ended up with a weird fixation with assuming everything was some sort of approximately geometric sequence and would always conclude none of the options matched. Before I gave it the eval_js tool it still did a pretty good job making choices, but the explanations were worse.
- lelandfe 2y agoUnaccountably correct is not a frequent quality I see in humans
- skywhopper 2y agoThis is because LLMs do not reason. They pattern-fit. The fact that that solves a lot of things humans often use reason to solve most likely speaks to the training data or unrecognized patterns in standardized tests, not to LLMs reasoning capability, which does not exist. To excuse their assumption of reasoning capabilities, the author in the FAQ snarkily points to “research” indicating evidence of reasoning—all of which was written by OpenAI and Microsoft employees who would not be allowed to publish anything to the contrary. It’s a shame people continue to buy into the hype cycle on new tech. Here’s a hint: if the creators of VC-backed tech make extraordinary claims about it, you should assume it’s heavily exaggerated if not an outright lie.
- quantadev 2y agoLLMs can't explain their past behavior because any time you ask them to, all they're really doing is reading the previous context from a previous inference and looking for any reason they can think of that they might have said that, so to speak. That is, they have absolutely no genuine recollection of what they were thinking at the time they said something in the past. Even with "Tree of Thought" approaches all you're doing is recording past conversations, states, and contexts, and your new inference asking for the "justification" of that, will be a similarly totally fake justification, because as I said they have no memory but only context. In my own app I can switch to a different LLM right in the middle of a conversation and the new LLM will just continue to always think it said everything in the prior context even though that's not the case.
- Terretta 2y ago> the new LLM will just continue to always think Even when you know, as you do, it's tough to avoid such characterizations.
- 01HNNWZ0MV43FF 2y agoIt might be fair to say they think but they can't directly introspect the thinking process. They can only confabulate reasoning post-hoc, or use science to try to understand themselves from the outside. Same as humans.
- stackghost 2y ago>Same as humans Hard disagree. LLMs are very accurately described by the "stochastic parrot" analogy that gets thrown around a lot. They do not "think" like humans at all, even if we use the word "think" because it's convenient.
- Der_Einzige 2y agoThe more tokens spent on getting a result, the more likely the result is to be accurate. “If it walks like a duck, quacks like a duck” etc
- mandevil 2y agoAlmost a decade ago I ran into a couple of problems where our logs started to change qualitatively on a SaaS tool I was supporting, but not in a way that printed more error level messages in logs. When someone complained about incorrect behavior in our code we could see clearly in the logs that the wrong code paths were being engaged, and when that started- people had carefully logged everything that was happening! They just hadn't logged it at the "Error" level so we never got alerts about it. And it would have been impracticable to have that level of alerting, because the message wasn't an error, it was just that due to other code changes we were now going down the wrong path for this span. So for a hackweek I built a tool to tokenize all of our log messages, and then grabbed all of our logs and built a gigantic n-dimensional vector for every five minute chunk of two days of those logs, then calculated the pythagorean difference for each of those five minute chunks, and looked at the biggest differences, most outlier five minute chunks. And they were all from 8-8:30AM CET on the two days (our company and most of our customers were US based, I just was looking at what timezones matched up to the interesting time). I said "okay, this looks interesting, let me see what is happening in the logs then" but it was impossible to figure out what the statistics were seeing. Because the math thinks in ways that human brains don't- it views the entire dataset simultaneously, and human brains just can't keep five minutes of busy log files in their working memory, but humans build narratives and the math can't understand that. So I ended up getting frustrated and giving up on the project. Because explaining in terms that I could understand and start debugging was the whole point of the project!
- jonathanyc 2y agoWow, I'm actually really interested in how you did that. Did the differences between chunks correspond to changes in word/token frequencies? Or did you embed the logs into some other space first? This is a pretty nifty idea; the only analyses I've ran on production logs were just simple SQL queries.
- mandevil 2y agoI wrote a python script that parsed all of our back-end code, looking for every log message, and used that as the basis for a giant n-dimensional vector representing every possible log message. Then I looked at every five minute chunk of our logs, and incremented the appropriate row in the vector for every message printed, then I used a clustering algorithm (don't remember which one, sorry, but it something from PyTorch I'm pretty sure, not at this company any more so I don't have the code any more) to identify the five vectors that were farthest from any other vector. I first computed the n-dimensional pythag difference between each pair of five minute chunks and selected the vectors with the highest combined difference but the true clustering algorithm was a cleaner choice. I did throw away a couple of messages, basically all the traffic from our uptime and health checks I threw away because they seemed like they would distort the data (if our health-checker went down that was the health-checkers fault, and we ought to get an actual alarm, having this alert would just be a duplicate alarm). Dunno, that might have been a mistake. It was a hackweek project, I'm not saying this was perfect- in fact, as I said, it never provided anything useful because it couldn't be explained! In our actual log files on the two days I looked at, all five of the outlier five minute chunks were from 8-8:30 AM Central European Time (our logs were in UTC, I just looked at different time zones and that seemed the most likely source of interesting behavior). And then when I looked at those times in the logs- and also when I eyeballed the ~1000 dimensional vectors for them- I couldn't tell what the clustering algorithm was seeing, because it was 'thinking' so differently from how I do. It wasn't like one of the rows in the vector was suddenly 150 and then went to 0 outside that half-an-hour, it was hard to see any patterns, So I couldn't set this up to alerts or anything, because it would have produced a whole lot of wild goose chases without a lot of further refinement. Talking with someone more expert in ML than I, he recommended trying to fine-tune a LLM to predict the next word (or maybe even the next log message depending on windows and verbosity of log messages) and potentially alerting when the differences between predicted and actual got too large. Maybe that would work, I dunno. But that was what I would have tried next hackweek, if I hadn't moved on from that company.
- distortionfield 2y agoThe house number you saw is actually part of the Lucas sequence. it’s related to and is a good approximation at large numbers of the Fibonacci but at small numbers becomes distorted. https://en.m.wikipedia.org/wiki/Lucas_number https://en.m.wikipedia.org/wiki/Lucas_number
- jonathanyc 2y agoHa, this is neat!! I’d never heard of the Lucas sequence before, thanks for the new knowledge!
- anon291 2y agoIn some ways it's obvious why. An LLM produces a probability distribution and then a random word is sampled. Imagine if you said to a secretary that you're 60% yes and 40% no, and she arbitrarily decided to write NO in your report and then a day later the board asked you why you made that decision. You'd be confused too.