2 ms·
Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of
by aftbit 17d ago
Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?