5 ms·
>It’s that it didn’t obey what it was told. I find you basically have to stop thinking of LLMs as software and start thinking of them as unpredictable animals.
by koboll 3y ago
>It’s that it didn’t obey what it was told.
I find you basically have to stop thinking of LLMs as software and start thinking of them as unpredictable animals. If you issue a command and expect strict obedience every time, you've already failed. Strict orders are really a tool to persuade certain behavior rather than some sort of reliable guardrail.
- moffkalast 3y agoSo the correct way to configure LLMs is to look at them sternly and yell "BAD DOG!" when they don't follow instructions and give them treats when they do?
- williamdclt 3y agoThe way to “configure” LLMs is training, yes!
- sebzim4500 3y agoThe technical term is Reinforcement Learning from Human Feedback (RLHF) but yes, that's basically what you do.
- moffkalast 3y agoHa I suppose it's exactly that.