4 ms·
A lot of people are worried about aligning superintelligent, self-improving AI. But I think it will be easier than aligning current AI, for the same reason that
by trott 2y ago
A lot of people are worried about aligning superintelligent, self-improving AI. But I think it will be easier than aligning current AI, for the same reason that it's easier to explain what you want to a human than it is to train a dog.
I posted my specific proposal here: https://olegtrott.substack.com/p/recursion-in-ai-is-scary-but-lets https://olegtrott.substack.com/p/recursion-in-ai-is-scary-bu...
Unlike previous ideas, it's implementable (once we have AGI-level language models) and works around the fact that much data on the Internet is false. I should probably call it Volition Extrapolated by Truthful Language Models.
- greenthrow 2y agoJust because you explain what you want to a human that doesn't mean they agree or will comply. That said, I see no reason to believe we are even on a path that can create AGI. LLMs don't actually understand or reason about anything.
- trott 2y ago> Just because you explain what you want to a human that doesn't mean they agree or will comply. Humans have innate desires that may conflict with the desires of other humans. A language model just looks for ways to continue texts. In doing so, it attempts to extrapolate what humans who authored the training data thought (If we ignore fabrications, which my approach proposes to address also) > LLMs don't actually understand or reason about anything. Current Transformer-based, SGD-trained language models fall short of AGI. But better algorithms can change that. And there is no reason to think that you'll see any warning signs in advance. No one said "I'll invent a faster Fourier transform algorithm in 5 years. Prepare yourselves."