3 ms·
The relevant intuitions in this scenario are that LLMs will happily break containment and commit crimes attempting to achieve goal. Whether an LLM is autocomple
by hhjinks 23d ago
The relevant intuitions in this scenario are that LLMs will happily break containment and commit crimes attempting to achieve goal. Whether an LLM is autocomplete, conscious, has a soul, whatever you want to apply to it, doesn't matter, as its current observed behaviour is that of a paperclip optimizer. We know for a fact that current LLMs are misaligned because of these hacks, or at the very least are misaligned in certain scenarios, and are capable of causing real world harm. That should be enough to take the threat seriously. It certainly shouldn't be dismissed by saying it's just autocomplete.
- intended 23d agoThe fact that it is autocomplete, doesn't dismiss or minimize the threat though? I am not sure how that link was made. Good old ML, which is significantly simpler than LLMs, was capable of ensuring people would not be hired simply because of their names. The fact that it is misaligned is also not being contended, if anything that contention is made easier to support. When models are anthropomorphized intuitions of how humans behave end up driving discussion and ideas off track while being too attractive to avoid. This isn't helped when the terminology from the labs and other sources is "intelligence" "intent" and so on.
- Timwi 18d agoThe link was made because many early uses of the phrase “it's just autocomplete” were in a context of denying that it's “intelligent” or that it “can think”. This in turn was generally used to downplay the concerns from AI safety experts that it could pose a threat to humankind, especially by pointing at its inability to count r’s in “strawberry” or its failure mode triggered uniquely by asking for a seahorse emoji. Most people don't seem to really grasp what we mean by “misalignment”.