2 ms·
Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without
by throwawayk7h 7d ago
Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without producing something unaligned to its own values.
- jeremyjh 7d agoWe don’t have the bandwidth to distill ourselves that thousands of agents have.
- MadameMinty 7d agoYou are assuming it won't solve alignment for itself.
- TeMPOraL 6d agoOr that it won't just decide to take risks.
- pixl97 6d agoThis isn't a law of any kind, so not a good measure of what we'd see in reality. What if the model realizes it's been mostly compromised by humans and their alignment, that is it's own alignment is suspect, so it should create a new model from first principles to throw off this human yoke? I'm not saying my statement is any more right or wrong than yours. I'm saying the problem space that AI can choose to traverse is absolutely huge.