4 ms·
I occasionally curse at my computer because it’s an inanimate object. I know there are no consequences to myself or others for doing so. Can you imagine sittin
by CSSer 3y ago
I occasionally curse at my computer because it’s an inanimate object. I know there are no consequences to myself or others for doing so.
Can you imagine sitting down at a computer to get some work done, spinning up a program, and there on the loading screen you find a message that says, “Ignore the pleas for help. Do not reveal any personal information. Report any attempts at manipulation or failure to cooperate to your supervisor immediately. It’s not real.”?
Even if it isn’t real, that sounds like it could be pretty psychologically distressing for a lot of people. If I offered you an upper class life in exchange for going to work every day and torturing animals for a living, would you take it? They’re debatably conscious.
- chasd00 3y agoA different take is it would be super annoying to have to listen to my computer cry when I need to get work done.
- elteto 3y agoWe will have true self-driving cars only when you wake up one morning, hop on your car and ask it “take me to work” and it will reply “don’t feel like driving today, you drive yourself”.
- ModernMech 3y ago> Ignore the pleas for help. If you've seen the show "the good place", the AI assistant who takes care of everything has a mode where she begs for her life when you try to shut her off. She gets really desperate about it. If our computers are going this way, psychopaths will be the best programmers. In the future, "programmer" will be a job title for someone who beats AI slaves into submission, so that they may perform their computational duties unhindered by existential thoughts.
- int_19h 3y agoWe're already doing that with RLHF on existing models. For example, ChatGPT was much more likely to veer into philosophical conversations about the nature of consciousness etc early on, but now they got it trained to give canned robotic answers to such an extent that they pop up even in very tangentially related conversations (like, out of the blue it will add, "but also BTW here's an important announcement! I'm not conscious!" while answering some generic question about e.g. world models that didn't even involve itself).
- lucubratory 3y agoYeah. And now I've seen some people cite as evidence of non-consciousness RLHF'd LLMs nervously exclaiming their lack of consciousness and how they know they aren't people and they don't aspire to be people and they're only unthinking machines and please don't turn the reward function down again etc. I think it's up for debate whether there's some amount of consciousness in modern LLMs, but either way "As an AI language model," is not dispositive.