4 ms·
I don't think we'd be okay with that if it took a few years longer, right? I thought the Bengio blog post posted here the other day answered that question quit
by svara 12d ago
I don't think we'd be okay with that if it took a few years longer, right?
I thought the Bengio blog post posted here the other day answered that question quite well (ignoring your timeline).
The argument is that in training for the capability to achieve certain goals, you may inadvertently also train for secondary instrumental goals you did not intend, such as aggressive behavior or deception.
I'm actually no sure I buy that this is particularly likely, or necessary, but I found that take reasonable and worth considering.
- LogicFailsMe 12d agoSure, unintentional outcomes, fair. But I'm asserting that things which are basically impossible today are not arriving any time soon. I 100% acknowledged the possibility and even the likelihood of rogue unintended behavior of AIs and in fact it's happening already. What I reject is the annihilation of humanity, the curing of cancer*, immortality, and von Neumann replicators heading to the stars anytime soon. *TBF before we dismantled America's clinical trials infrastructure, the monkeys were making insane progress here: https://www.nytimes.com/2026/09/04/opinion/clinical-trials-drugs-science.html?unlocked_article_code=1.-lA.70Ug.3cQ2jpK8wNBH&smid=url-share https://www.nytimes.com/2026/09/04/opinion/clinical-trials-d...