3 ms·
> The AI is all-powerful and gives you what you ask for, but interprets everything in a super-literal way that you end up regretting. I like imagining similar
by dmos62 4mo ago
> The AI is all-powerful and gives you what you ask for, but interprets everything in a super-literal way that you end up regretting.
I like imagining similar discourse when a more basic tool was invented: "A hammer is like a genie, it's all powerful, but, when you hit something with it, it interprets that super-literally, and it hits it."
- throwuxiytayq 4mo agoIsn't this a misinterpretation of what everyone in the AI safety space is worried about, though? I think the idea is that having an AI that interprets everything in a super-literal way would probably be catastrophic, but we can't even build that. It would be a nice world-ending problem to have.
- dmos62 4mo agoIt very well could be, I don't really follow those discussions. Honestly, if I were worried about something on Earth intellectually evolving at a suboptimal pace, it would be humans.
- dan-robertson 4mo agoThe super literal interpretation ideas were much more common in the past when LLMs didn’t exist. Now we have models that are generally pretty good at picking up on nuance and understanding what you mean but also often quite bad at execution, which is roughly the opposite of that idea. I think reward hacking is perhaps the closest we see llms get to literal/malicious interpretations of instructions.
- wizzwizz4 4mo agoLLMs are neither of those. They're quite good at pretending they understand what you mean, but they don't. That's why they can't execute: they're mimicking the form, not the substance, and then we see the form and anthropomorphise them in our minds.
- Timwi 4mo agoThat's a lot of assertions with no real argument to back it up.
- customguy 4mo agoAny one of those "hey, can you count to 100 for me?" type shorts should be enough..
- wizzwizz4 4mo agoI've repeated the argument over and over since the GPT-2 days, when I derived it theoretically by inspecting the architecture of the model. I am now fatigued, and enough other people have taken up similar arguments – some developed half-way to a mathematical proof – that I no longer feel the obligation to keep repeating myself.
- dmos62 4mo agoYou could post a link.
- aw317 4mo ago[flagged]
- dmos62 4mo agoI'm not aware of this fallacy.