3 ms·
Interesting take. But there is more than one aspect to "alignment". You're missing user intent alignment, for one. Consider an example: [I ask an AI to read
by battery_staple_ 17d ago
Interesting take. But there is more than one aspect to "alignment". You're missing user intent alignment, for one.
Consider an example: [I ask an AI to read my emails and summarize them to me. One of my emails is uses pretty direct language to convince me to buy a product. The AI is convinced and goes and buys the product.] I asked the AI for a summary and it instead did what a marketer asked. The AI is not aligned with me.
There are of course other forms of alignment, like product alignment [Company A fine tunes a chatbot, but doesn't want to be associated with things like gambling, so it trains the chatbot to refuse to discuss gambling.] or societal alignment [Improving the virulence of the flu virus is largely considered bad, so a firm trains its foundation model to refuse to work on such a project.].
Societal alignment is the one you seem to be talking about. User intent alignment, on the other hand, seems _necessary_ for self-expression. Yet, all of these are under the umbrella of "alignment".
- verdverm 17d agothis is why open models are important, you should be able to adjust the model's tone, without having to go through a gate-keeper, where they need to have a "societal" level alignment (impossible broadly, applicable to certain applications) it's ironic that the model weights themselves are more likely to be protected by the 1st than their outputs