4 ms·
I think this is interesting enough for a post in and of itself: https://arxiv.org/abs/2505.22954 https://arxiv.org/abs/2505.22954
by datameta 1y ago
I think this is interesting enough for a post in and of itself: https://arxiv.org/abs/2505.22954 https://arxiv.org/abs/2505.22954
- wiz21c 1y agoFrom the article abstract: "All experiments were done with safety precautions (e.g., sandboxing, human oversight)." Do the authors really believe "safety" is necessary, that is, there is a risk that somethign goes wrong ? What kind of risk ?
- datameta 1y agoFrom what I understand, alignment and interpretability were rewarded as part of the optimization function. I think it is prudent that we bake in these "guardrails" early on.