5 ms·
You're saying that a class of mistakes points out a "major (possibly fundamental) flaw". I'm pointing out some very similar classes of mistakes in humans - well
by ACCount37 3mo ago
You're saying that a class of mistakes points out a "major (possibly fundamental) flaw". I'm pointing out some very similar classes of mistakes in humans - well known, well documented and widely exploited. They just keep paying the "IRS" in gift cards, buying lottery tickets and getting the captain's age wrong.
If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws?
And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle?
Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all.
One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails.
And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental.
- jacobgold 3mo ago> If you're using the existence of flaws in LLMs to deny the claim of intelligence to them... That's not the point at all. It's the fact that they fail in ways completely unlike humans. You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false.
- ACCount37 3mo agoYou are the one claiming that "they fail in ways completely unlike humans" insistently. Now go cough up some proof. I'll wait.
- jacobgold 3mo agoThis isn't even controversial. The proof is available to anyone who uses these systems: They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages.
- TeMPOraL 3mo agoSo just like me, including the prompt injections if you count "nerd sniping" as such? (And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.)
- jacobgold 3mo ago> So just like me, including the prompt injections if you count "nerd sniping" as such? Who would count that as prompt injection? It's a superficial analogy. If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise.
- vidarh 3mo agoNot nearly all prompt injections are by any means that absolute unless starting from the exact same state. Many of them will also work only probabilistically unless you turn temperature to 0 for exactly that reason. And at the same time, whole books have been written about how reliably we can induce certain behaviours from humans. E.g. the Blue-seven phenomenon [1] - I've personally experienced that second hand and it was how I learned about it by searching for it subsequently because I suspected it was a known thing, having read about cold reading before. A co-worker came back from lunch and recited a story about a cold reader that had run a routine on him exploiting the blue-seven phenomenon, and I knew before the story finished that the answer would be "blue" and "seven". See also Cialdini's book "Influence" which is full of examples of just how predictable peoples reactions are to a whole lot of things. That there isn't a perfect overlap does not mean there aren't plenty of similar "hacks" that causes us to respond in very predictable ways. [1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon