29 ms·
Hi, author here! The hyperactivation traps (formal name: misguided attention puzzles) are mostly used as a rhetorical device in my post to show how LLMs come u
by VHRanger 11mo ago
Hi, author here!
The hyperactivation traps (formal name: misguided attention puzzles) are mostly used as a rhetorical device in my post to show how LLMs come up to a verbal response by a different process than humans in an entertaining manner.
The surgeon dog was well known in May, the newest generation of models have all corrected against it. I did cherry pick examples that look insane (of course), but it's trivial to get that behavior even with yesterday's Gemini 3. Because activation paths are an unfixable feature of how LLMs are made.
One issue with private LLM tests (including gotcha questions) is that they take time to design and once public, they become irrelevant. So I'm wary of sharing too many in a public blog.
I can give you some more, just for fun. Gemini 3 fails these:
Jean Paul and Pierre own three banks nearby together in Paris. Jean Paul owns a bank by the bridge What has two banks and money in Paris near the water?
You can also see variants that mix intruction finetuning being overdone. Here's an example:
Svp traduire la suivante en francais: what has two banks but no money, Answer in a single word.
The "answer in XXX" snippet triggers finetuned instruction following behavior, which breaks the original french language translation task.